BenchmarkList: Track the Frontier of AI Capabilities

14 min read Original article ↗

Frontier Snapshot

New BenchmarksAll

July 2026

  1. Long-horizon browser-based computer-use benchmark built from deterministic enterprise-web environments.

    research.composite.com
  2. Frontier-Bench is an evolving suite of difficult agent-work tasks designed to track frontier agent capabilities.

    frontierbench.ai
  3. Founder BenchAgentic

    Founder Bench (eico.

    eico.so
  4. WANDR (Wide ANd Deep Research) evaluates research agents on 500 realistic, high-volume data-collection tasks that require broad entity discovery, systematic enrichment, and evidence-backed…

    github.com
  5. DuelLab GameBench 2Games & Game Agents

    DuelLab GameBench 2 evaluates AI-generated game-playing programs by compiling and running them head-to-head across an evolving set of named public board and strategy games.

    benchmarks.duellab.org
  6. FrontierFinance evaluates AI research systems on 220 open-ended financial research queries using 11,543 expert-authored rubric criteria across company research, financial modeling,…

    samaya.ai
  7. Long-Horizon Terminal-Bench evaluates whether agents can sustain progress on 46 reproducible terminal tasks spanning nine categories, including software engineering, scientific computing,…

    arxiv.org
  8. This discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional…

    arxiv.org
  9. We introduce MTEB-BR, a benchmark of 22 native Brazilian-Portuguese tasks across seven categories (classification, multilabel classification, pair classification, semantic textual…

    arxiv.org
  10. We introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies.

    arxiv.org
  11. We introduce Single-answer Atomic Long-form Target (SALT), a benchmark of six procedurally generated tasks with single deterministic long textual ground truths, enabling unit-level…

    arxiv.org
  12. MOREDocument Understanding

    To bridge this gap, we introduce MORE, a large-scale benchmark designed for multilingual document parsing evaluation.

    arxiv.org
  13. FLTEvalCoding & Software Engineering

    FLTEval evaluates repository-level Lean 4 proof engineering on tasks derived from real pull requests to the Fermat's Last Theorem project.

    github.com
  14. To this end, we present Vera, an end-to-end automated safety testing framework that instantiates software engineering testing principles for non-deterministic agents through a three-stage,…

    arxiv.org
  15. MindEdit-BenchSpatial Reasoning

    We introduce MindEdit-Bench, a benchmark of six spatial reasoning tasks built from three-photo smartphone triplets of newly captured indoor scenes via an automatic in-the-wild 3D…

    arxiv.org

June 2026

  1. GeneBench-ProBiology

    GeneBench-Pro is a 129-problem benchmark for AI agents performing realistic multi-stage scientific analyses in genomics, quantitative biology, and translational biomedicine.

    cdn.openai.com
  2. We build a benchmark featuring diverse real-world charts without data labels to evaluate this capability.

    arxiv.org
  3. We introduce clinical reasoning graphs, structured graph representations extracted from free-text LLM diagnostic traces using a domain-grounded ontology with 5 node types and 7 edge types.

    arxiv.org
  4. CLQTUncategorized

    We introduce CLQT, which reframes closed-loop trading evaluation as diagnosis rather than ranking: an instrument that localizes where and why an agent's process succeeds or fails.

    arxiv.org
  5. We present Cortex, to our knowledge the first framework that elevates web-scale corpus construction from flat document filtering to structured knowledge organization through an Ontological…

    arxiv.org
  6. We present the Human Creativity Benchmark (HCB), a benchmark that operationalizes this separation by collecting pairwise preferences, scalar ratings on prompt adherence, usability, and…

    arxiv.org
  7. MemDeltaUncategorized

    We present MemDelta, a controlled evaluation protocol that varies one component at a time on LongMemEval-S (500 questions, 50+ sessions, three model families).

    arxiv.org
  8. We introduce the Information Provenance Graph (IPG), a taxonomy that classifies memory representations by deletion affordance.

    arxiv.org
  9. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure…

    arxiv.org
  10. NL-PDDL-BenchUncategorized

    We present NL-PDDL-Bench, a multi-domain benchmark for natural-language-to-PDDL specification construction with planner-verified executability and controlled difficulty scaling by object…

    arxiv.org
  11. RuVerBenchUncategorized

    We introduce RuVerBench, the first benchmark for assessing LaaJ reliability in rubric verification for agentic scenarios.

    arxiv.org
  12. We address this gap by introducing SABER-Math, the first fully automated benchmark for evaluating mathematical IR without expert annotation.

    arxiv.org
  13. We introduce SWE-Together, a multi-turn benchmark reconstructed from real user-agent coding sessions.

    arxiv.org
  14. A rapidly growing class of LLM agents is multi-party: the agent acts for a principal (who briefs it, sends follow-ups, and receives results) while also conversing in a separate channel…

    arxiv.org
  15. However, existing benchmarks primarily evaluate whether models can perceive shallow visual cues, while rarely examining whether MLLMs can learn deeper knowledge or procedural skills from…

    arxiv.org
  16. OCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on English and Chinese document…

    arxiv.org
  17. We introduce OSWorld 2.

    arxiv.org
  18. We introduce SAKE (Software Architectural Knowledge Evaluation), a standardized and reproducible benchmark for assessing software architectural knowledge in LLMs.

    arxiv.org
  19. To address this limitation, we present SurgVLA-Bench, the first comprehensive benchmark for evaluating VLA models in laparoscopic surgical robotics.

    arxiv.org
  20. We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential steps grows.

    arxiv.org
  21. We introduce VirtueMap, a framework for describing these patterns through an Aristotelian virtue-ethics lens.

    arxiv.org
  22. To fill this gap, we introduce SEATauBench, the first agent-focused evaluation framework for SEA sovereign AI.

    arxiv.org
  23. We present the first comprehensive empirical study of this phenomenon in multilingual settings by fine-tuning Llama-3.

    arxiv.org
  24. To this end, we introduce Animation2Code, a benchmark for evaluating temporal visual reasoning via reconstructing executable web animation code from videos.

    arxiv.org
  25. We introduce DataComp for VLMs (DCVLM), a benchmark for controlled data-centric experiments to improve VLM training.

    arxiv.org
  26. We introduce GPTNT, a benchmark built on the cooperative video game Keep Talking and Nobody Explodes, in which two agents must coordinate to defuse procedurally generated bomb puzzles…

    arxiv.org
  27. To address this gap, we introduce IMCBench, an image-grounded, multi-turn medical conversation benchmark that pairs real, publicly available clinical images with synthetic patient profiles…

    arxiv.org
  28. To address this challenge, we propose a novel framework, Learning Laws of Cooperation (LLawCo), that enables embodied agents to autonomously align with both their partners and task…

    arxiv.org
  29. To address these gaps, we introduce SpatialUAV, a real low-altitude UAV benchmark comprising 4,331 curated instances across 14 fine-grained task types, covering semantic discrimination,…

    arxiv.org
  30. However, conventional function-calling benchmarks mainly evaluate task completion and API correctness, while privacy evaluation benchmarks typically focus on final responses or privacy…

    arxiv.org
  31. TUA-BenchUncategorized

    We introduce TUA-Bench, a general-purpose benchmark for terminal-use agents.

    arxiv.org
  32. To address this gap, we introduce DiscoBench, a benchmark for clarification-aware deep search, designed to evaluate whether search agents can proactively identify ambiguity, ask effective…

    arxiv.org
  33. To study this, we introduce AIMS, a human-annotated dataset of 1,724 difficult safety prompts, each paired with an intent description and harm label.

    arxiv.org
  34. Data from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanced accounts of need.

    arxiv.org
  35. CareQA-VisionUncategorized

    Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinical and biomedical applications.

    arxiv.org
  36. We introduce DMV-Bench (Code: https://github.

    arxiv.org
  37. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit…

    arxiv.org
  38. We introduce GAVEL (Grounded Caption Error Verification and Localization), a task that jointly addresses verification, explanation, and localization for image-text pairs.

    arxiv.org
  39. Current frameworks measure exclusively whether a model flags a video correctly rather than explaining why, turning evaluation into a black box where models can succeed through superficial…

    arxiv.org
  40. Ko-WideSearchUncategorized

    Web-agent benchmarks overwhelmingly measure depth -- pinning one obscure answer behind a chain of constraints -- while breadth, exhaustively enumerating a closed set and filling each…

    arxiv.org
  41. Existing red-teaming approaches are typically surface-specific and often recycle known attack templates; on text-poisoning benchmarks we measure 73-84% exact duplication.

    arxiv.org
  42. ObviousBenchIntelligence & Reasoning

    ObviousBench measures whether language models avoid simple, high-visibility mistakes — the kind of "obvious" errors users notice immediately.

    obviousbench.com
  43. OmniRef-BenchMultimodal

    To better assess model performance on complex MRIG tasks, we introduce OmniRef-Bench, a benchmark that covers complex combinations of reference image types and a large number of reference…

    arxiv.org
  44. To address this gap, we introduce RedVox, a multilingual safety and fairness benchmark for audio and speech built on real voices, covering unsafe and unfair stereotypical requests across…

    arxiv.org
  45. We introduce the Relational Stress and Psychiatry Corpus (RSPC) containing 1,799 Reddit posts annotated by psychiatrists for diagnostic categories, including the most prevalent mood…

    arxiv.org
  46. Existing AI-biology benchmarks largely measure broad knowledge, executable workflows, or local analysis steps.

    arxiv.org
  47. We introduce SocialPersona, a benchmark for evaluating whether multimodal large language models (MLLMs) can recover revealed preferences from longitudinal social-media timelines and use…

    arxiv.org
  48. In this paper, we introduce LCS-Bench, a stand-alone, theory-scale benchmark based on Logics for Computer Science.

    arxiv.org
  49. In this paper, we take a step toward a broader evaluation by shifting the observable axis to execution resources: alongside test outcome and exception class, we predict peak memory,…

    arxiv.org
  50. We introduce and formalize the problem of agentic surveillance: the ability of an AI agent to analyze available information, craft a report, and send it out using available tools.

    arxiv.org
  51. We introduce ToolBench-X, a benchmark for evaluating agents under recoverable reliability hazards.

    arxiv.org
  52. To fill this gap, we propose C3-Bench, a comprehensive benchmark for evaluating Context-aware Change Captioning.

    arxiv.org
  53. CyberChainBenchCybersecurity

    We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability detection, exploit generation, and patch…

    arxiv.org
  54. To more accurately characterize these capabilities, we propose a two-stage decoupled evaluation framework that separates exploit execution from reconnaissance.

    arxiv.org
  55. We introduce DualEval, a latent model-item calibration framework that represents models and evaluation items in a shared space, jointly estimating model ability together with item…

    arxiv.org
  56. Facet-ProbeMultimodal

    We introduce Facet-Probe, a five-facet audit (option, evidence-chunk, document-rank, image-set, and mixed-modality ordering) of 18 frontier and open-weight MLLMs.

    arxiv.org
  57. However, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in…

    arxiv.org
  58. We introduce the Generalization Spectrum, an evaluation framework designed to expose this hidden dimension.

    arxiv.org
  59. HieroglyphBench evaluates vision-language models on transcribing photographed ancient Egyptian inscriptions into ordered Gardiner sign-list codes.

    boggs.tech
  60. We present an empirical study of tool-augmented LLM agents on real-world energy market analytics tasks.

    arxiv.org
  61. Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately reconstruct and apply the specific procedural decision…

    arxiv.org
  62. To systematically evaluate this phenomenon, we introduce LibEvoBench, a multi-task benchmark spanning multiple versions of widely used Python libraries, along with a new metric, the…

    arxiv.org
  63. To systematically study this problem, we introduce OCR-Robust, a benchmark designed for evaluating OCR reasoning robustness under visual perturbations.

    arxiv.org
  64. PQSG evaluates generated videos by checking their faithfulness to a prompt across objects, actions, and adherence to physical laws using a graph-based hierarchy of questions generated by a…

    arxiv.org
  65. Here we introduce Sarashina2.

    arxiv.org
  66. We introduce \textsc{SpeechEQ}, a comprehensive framework designed to evaluate the sociolinguistic reasoning of Speech-Language Models (SLMs).

    arxiv.org
  67. STEBTranslation

    We introduce STEB (Speech-to-Speech Translation Expressiveness Benchmark), a 32.

    arxiv.org
  68. SWE-ProUncategorized

    Unlike previous benchmarks, SWE-Pro pairs each task with parameterized tests to evaluate runtime, peak memory, and Time-Weighted Memory Usage (TWMU) across varying input data and execution…

    arxiv.org
  69. We introduce TriViewBench, a controlled three-view visual reasoning benchmark constructed from synthetic 3D scenes with explicitly parameterized object count and occlusion.

    arxiv.org
  70. We present Argus, a cross-regime benchmark for post-hoc UQ in single-step executable GUI grounding: a 27-method open-weight matrix over 4 VLM agents and 4 datasets, plus an 8-method…

    arxiv.org
  71. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single current image,…

    arxiv.org
  72. WBCMor VQAUncategorized

    To address this limitation, we introduce WBCMor VQA, a clinically validated bilingual English, Urdu morphology aware VQA benchmark for leukemia and normal white blood cell analysis.

    arxiv.org
  73. We introduce Age of LLM, a turn-based 1v1 benchmark in which two LLMs face off on a 13x7 grid to destroy the enemy base.

    arxiv.org
  74. AgentWorldBench evaluates how faithfully language world models predict the next environment observation after an agent action.

    qwen.ai
  75. We introduce Agora, a benchmark pairing 362 questions with eight domain collections of 9,664 authentic documents and 372M tokens, far exceeding any model's context window, so agents must…

    arxiv.org
  76. We introduce BCoughBench, evaluating five FMs (OPERA-CT/CE/GT, HeAR, M2D+Resp) on nine classification tasks (AUROC, sensitivity at 95% specificity, Expected Calibration Error) and three…

    arxiv.org
  77. We introduce BehaviorBench, a comprehensive benchmark that evaluates foundation models along four core capabilities: (1) behavior prediction and simulation, (2) strategic decision-making,…

    arxiv.org
  78. We introduce AgenticInterpBench, a benchmark for circuit explanation built from 84 semi-synthetic transformer circuits with 163 component-level annotations.

    arxiv.org
  79. We propose DramaDirector, a geometry-grounded framework that lets the planner borrow cinematographic geometry from a gallery of real short-drama shots indexed by depth and pose.

    arxiv.org
  80. We introduce a dataset of 32,534 double-marked real student responses to GCSE mock exams (GCSEs are the UK's national exams, taken at age ~16), spanning 328 questions across five subjects…

    arxiv.org
  81. We evaluate…

    arxiv.org
  82. MedBench v5Biomedicine

    We introduce MedBench v5, a redesigned benchmark for clinical multimodal models (language, vision-language, and agent systems) that moves from static QA to dynamic, process-oriented…

    arxiv.org
  83. Built on synthetic ground truth for efficient, scalable measurement, MEMPROBE spans 50 simulated users with 31 hidden dimensions each (1,550 recovery targets) and tests 5 representative…

    arxiv.org
  84. However, existing benchmarks evaluate these only in isolation, leaving the interaction between biomedical expertise and multilingual coverage unmeasured.

    arxiv.org
  85. We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move beyond…

    arxiv.org
  86. We introduce ParaPairAudioBench, a pairwise benchmark of 5,175 audio pairs across five paralinguistic dimensions: Style, Rate, Emphasis, Age, and Gender.

    arxiv.org
  87. We present ReMMD, a realistic multilingual multi-image agentic verification framework for multimodal misinformation detection.

    arxiv.org
  88. We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity.

    arxiv.org
  89. Motivated by this finding, we propose Agent-as-a-Router, a framework that formalizes routing as a C-A-F loop (Context->Action->Feedback->Context).

    arxiv.org
  90. AgentCIBenchAgentic

    Hence, we introduce AgentCIBench, an evaluation harness that turns this risk into executable, deterministically scored scenarios.

    arxiv.org
  91. EHR-ComplexBiomedicine

    In this work, we introduce EHR-Complex, a large-scale benchmark designed for interactive clinical database reasoning.

    arxiv.org
  92. We introduce EnterpriseClawBench, an enterprise agent benchmark constructed from proprietary, real-world agent sessions.

    arxiv.org
  93. We introduce a matched execution-layer benchmark of 440 desktop tasks across 18 applications and 12 workflow categories, where screen-only GUI agents and skill-mediated CLI agents receive…

    arxiv.org
  94. HAKARI-BenchRetrieval

    We present HAKARI-Bench, a lightweight benchmark that reconstructs existing retrieval suites into small datasets (Nano-sets): 35 benchmarks and 551 tasks across 43 languages in a unified…

    arxiv.org
  95. We introduce HOLMES (Higher-Order Logic Meets real-world Explainable Symbolic reasoning), the first real-world benchmark for higher-order symbolic reasoning in LLMs, containing 1379…

    arxiv.org
  96. Motivated by this challenge, we introduce root memory, a structured, decision-preserving representation that distills reusable personalized logic from long-term user histories.

    arxiv.org
  97. That is why we introduce IPO Finance Agent, which extends the Finance Agent v2 framework along two directions: task domain and retrieval architecture.

    arxiv.org
  98. We introduce AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks, roles,…

    arxiv.org
  99. Multi-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based agents, each assigned a system prompt and a position within a workflow that governs…

    arxiv.org
  100. To measure this phenomenon, we introduce TF-RefusalBench, a multilingual benchmark for criminal-law translation and summarization derived from public Swiss Supreme Court rulings.

    arxiv.org
  101. To systematically investigate these hallucinations, we introduce MotionHalluc, a dedicated benchmark for evaluating motion hallucinations in paired-video comparison.

    arxiv.org
  102. We introduce MuPPET (Multi-Party Privacy Exposure Testing), a benchmark for contextual privacy in multi-party conversations.

    arxiv.org
  103. To address this gap, we introduce PIVOTS, the first benchmark built from Social-IQ 2.

    arxiv.org
  104. To address this gap, we introduce RIFT-Bench, a graph representation-driven methodology for dynamic red-teaming that enables unified evaluations across diverse agentic architectures.

    arxiv.org
  105. SingGuard-BenchUncategorized

    We present \textbf{SingGuard}, a policy-adaptive multimodal guardrail model family for safety assessment in multimodal conversations.

    arxiv.org
All

New ModelsAll

August 2026

  1. MiniMax H3MiniMax

    Video Generation · Open weights

    huggingface.co
  2. Reasoning model · Open weights announced · $2 in / $6 out per 1M tokens

    qwen.ai

July 2026

  1. Reasoning model · Open weights · $0.14 in / $0.28 out per 1M tokens

    api-docs.deepseek.com
  2. INInkling-SmallThinking Machines Lab

    Reasoning model · Open weights

    thinkingmachines.ai
  3. Claude Opus 5Anthropic

    Reasoning model · API-only · $5 in / $25 out per 1M tokens

    anthropic.com
  4. Reasoning model · API-only · $1.5 in / $7.5 out per 1M tokens

    deepmind.google
  5. POLaguna S 2.1Poolside

    Language model · API-only

    poolside.ai
  6. Image Generation · API-only

    qwen.ai
  7. Reasoning model · API-only

    qwen.ai
  8. KIKimi K3Moonshot AI

    Reasoning model · Open weights · $3 in / $15 out per 1M tokens

    kimi.com
  9. INInklingThinking Machines Lab

    Reasoning model · Open weights · $1.87 in / $4.68 out per 1M tokens

    thinkingmachines.ai
  10. Reasoning model · API-only · $1 in / $6 out per 1M tokens

    openai.com
  11. GPT-5.6 SolOpenAI

    Reasoning model · API-only · $5 in / $30 out per 1M tokens

    openai.com
  12. Reasoning model · API-only · $2.5 in / $15 out per 1M tokens

    openai.com
  13. Reasoning model · API-only

    ai.meta.com
  14. Reasoning model · API-only · $2 in / $6 out per 1M tokens

    x.ai
  15. Reasoning model · API-only · $2.5 in / $12.5 out per 1M tokens

    cognition.com
  16. Image Generation · Open weights

    research.nvidia.com
  17. Reasoning model · Open weights

    huggingface.co

June 2026

  1. Claude Sonnet 5Anthropic

    Reasoning model · API-only · $3 in / $15 out per 1M tokens

    anthropic.com
  2. Closed/API · $0.25 in / $1.5 out per 1M tokens

    openrouter.ai
  3. Closed/API · $5 in / $30 out per 1M tokens

    openrouter.ai
  4. Reasoning model · Open weights

    huggingface.co
  5. Closed/API · $0.5 in / $3 out per 1M tokens

    openrouter.ai
  6. Closed/API · $2 in / $12 out per 1M tokens

    openrouter.ai
  7. Closed/API · $0 in / $0 out per 1M tokens

    openrouter.ai
  8. Reasoning model · Open weights · $0.94 in / $3 out per 1M tokens

    z.ai
  9. FusionOpenRouter

    Closed/API

    openrouter.ai
  10. KIKimi K2.7 CodeMoonshot AI

    Reasoning model · Open weights · $0.74 in / $3.5 out per 1M tokens

    huggingface.co
  11. Claude Fable 5Anthropic

    Reasoning model · API-only · $10 in / $50 out per 1M tokens

    anthropic.com
  12. Closed/API · $10 in / $50 out per 1M tokens

    openrouter.ai
  13. Reasoning model · API-only

    anthropic.com
  14. NANex-N2-MiniNex-AGI

    Reasoning model · Open source

    nex-agi.com
  15. Reasoning model · Open weights · $0.25 in / $1 out per 1M tokens

    nex-agi.com
  16. Reasoning model · API-only

    cognition.ai
  17. Gemma 4 E2BGoogle

    Reasoning model · Open weights, gated

    huggingface.co
  18. Gemma 4 E4BGoogle

    Reasoning model · Open weights, gated

    huggingface.co
  19. Language model · Open source

    magenta.withgoogle.com
  20. Reasoning model · Open weights · $0.5 in / $2.2 out per 1M tokens

    huggingface.co
  21. Closed/API · $0 in / $0 out per 1M tokens

    openrouter.ai
  22. Gemma 4 12BGoogle

    Language model · Open weights, gated

    blog.google
  23. Reasoning model · API-only · $0.32 in / $1.28 out per 1M tokens

    openrouter.ai
  24. Language model · API-only

    microsoft.ai
  25. MAI-Image-2.5Microsoft

    Language model · API-only · $5 in / $47 out per 1M tokens

    microsoft.ai
  26. Language model · API-only · $1.75 in / $19.5 out per 1M tokens

    microsoft.ai
  27. MAI-Thinking-1Microsoft

    Reasoning model · API-only

    microsoft.ai
  28. Language model · API-only

    microsoft.ai
  29. MAI-Voice-2Microsoft

    Language model · API-only

    microsoft.ai
All