Frontier Snapshot
New BenchmarksAll
July 2026
Long-horizon browser-based computer-use benchmark built from deterministic enterprise-web environments.
research.composite.comFrontier-Bench is an evolving suite of difficult agent-work tasks designed to track frontier agent capabilities.
frontierbench.aiFounder BenchAgentic
Founder Bench (eico.
eico.soWANDR (Wide ANd Deep Research) evaluates research agents on 500 realistic, high-volume data-collection tasks that require broad entity discovery, systematic enrichment, and evidence-backed…
github.comDuelLab GameBench 2Games & Game Agents
DuelLab GameBench 2 evaluates AI-generated game-playing programs by compiling and running them head-to-head across an evolving set of named public board and strategy games.
benchmarks.duellab.orgFrontierFinance evaluates AI research systems on 220 open-ended financial research queries using 11,543 expert-authored rubric criteria across company research, financial modeling,…
samaya.aiLong-Horizon Terminal-Bench evaluates whether agents can sustain progress on 46 reproducible terminal tasks spanning nine categories, including software engineering, scientific computing,…
arxiv.orgThis discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional…
arxiv.orgWe introduce MTEB-BR, a benchmark of 22 native Brazilian-Portuguese tasks across seven categories (classification, multilabel classification, pair classification, semantic textual…
arxiv.orgWe introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies.
arxiv.orgWe introduce Single-answer Atomic Long-form Target (SALT), a benchmark of six procedurally generated tasks with single deterministic long textual ground truths, enabling unit-level…
arxiv.orgMOREDocument Understanding
To bridge this gap, we introduce MORE, a large-scale benchmark designed for multilingual document parsing evaluation.
arxiv.orgFLTEvalCoding & Software Engineering
FLTEval evaluates repository-level Lean 4 proof engineering on tasks derived from real pull requests to the Fermat's Last Theorem project.
github.comTo this end, we present Vera, an end-to-end automated safety testing framework that instantiates software engineering testing principles for non-deterministic agents through a three-stage,…
arxiv.orgMindEdit-BenchSpatial Reasoning
We introduce MindEdit-Bench, a benchmark of six spatial reasoning tasks built from three-photo smartphone triplets of newly captured indoor scenes via an automatic in-the-wild 3D…
arxiv.org
June 2026
GeneBench-ProBiology
GeneBench-Pro is a 129-problem benchmark for AI agents performing realistic multi-stage scientific analyses in genomics, quantitative biology, and translational biomedicine.
cdn.openai.comWe build a benchmark featuring diverse real-world charts without data labels to evaluate this capability.
arxiv.orgWe introduce clinical reasoning graphs, structured graph representations extracted from free-text LLM diagnostic traces using a domain-grounded ontology with 5 node types and 7 edge types.
arxiv.orgCLQTUncategorized
We introduce CLQT, which reframes closed-loop trading evaluation as diagnosis rather than ranking: an instrument that localizes where and why an agent's process succeeds or fails.
arxiv.orgWe present Cortex, to our knowledge the first framework that elevates web-scale corpus construction from flat document filtering to structured knowledge organization through an Ontological…
arxiv.orgWe present the Human Creativity Benchmark (HCB), a benchmark that operationalizes this separation by collecting pairwise preferences, scalar ratings on prompt adherence, usability, and…
arxiv.orgMemDeltaUncategorized
We present MemDelta, a controlled evaluation protocol that varies one component at a time on LongMemEval-S (500 questions, 50+ sessions, three model families).
arxiv.orgWe introduce the Information Provenance Graph (IPG), a taxonomy that classifies memory representations by deletion affordance.
arxiv.orgDespite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure…
arxiv.orgNL-PDDL-BenchUncategorized
We present NL-PDDL-Bench, a multi-domain benchmark for natural-language-to-PDDL specification construction with planner-verified executability and controlled difficulty scaling by object…
arxiv.orgRuVerBenchUncategorized
We introduce RuVerBench, the first benchmark for assessing LaaJ reliability in rubric verification for agentic scenarios.
arxiv.orgWe address this gap by introducing SABER-Math, the first fully automated benchmark for evaluating mathematical IR without expert annotation.
arxiv.orgWe introduce SWE-Together, a multi-turn benchmark reconstructed from real user-agent coding sessions.
arxiv.orgA rapidly growing class of LLM agents is multi-party: the agent acts for a principal (who briefs it, sends follow-ups, and receives results) while also conversing in a separate channel…
arxiv.orgHowever, existing benchmarks primarily evaluate whether models can perceive shallow visual cues, while rarely examining whether MLLMs can learn deeper knowledge or procedural skills from…
arxiv.orgOCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on English and Chinese document…
arxiv.orgWe introduce OSWorld 2.
arxiv.orgWe introduce SAKE (Software Architectural Knowledge Evaluation), a standardized and reproducible benchmark for assessing software architectural knowledge in LLMs.
arxiv.orgTo address this limitation, we present SurgVLA-Bench, the first comprehensive benchmark for evaluating VLA models in laparoscopic surgical robotics.
arxiv.orgWe introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential steps grows.
arxiv.orgWe introduce VirtueMap, a framework for describing these patterns through an Aristotelian virtue-ethics lens.
arxiv.orgTo fill this gap, we introduce SEATauBench, the first agent-focused evaluation framework for SEA sovereign AI.
arxiv.orgWe present the first comprehensive empirical study of this phenomenon in multilingual settings by fine-tuning Llama-3.
arxiv.orgTo this end, we introduce Animation2Code, a benchmark for evaluating temporal visual reasoning via reconstructing executable web animation code from videos.
arxiv.orgWe introduce DataComp for VLMs (DCVLM), a benchmark for controlled data-centric experiments to improve VLM training.
arxiv.orgWe introduce GPTNT, a benchmark built on the cooperative video game Keep Talking and Nobody Explodes, in which two agents must coordinate to defuse procedurally generated bomb puzzles…
arxiv.orgTo address this gap, we introduce IMCBench, an image-grounded, multi-turn medical conversation benchmark that pairs real, publicly available clinical images with synthetic patient profiles…
arxiv.orgTo address this challenge, we propose a novel framework, Learning Laws of Cooperation (LLawCo), that enables embodied agents to autonomously align with both their partners and task…
arxiv.orgTo address these gaps, we introduce SpatialUAV, a real low-altitude UAV benchmark comprising 4,331 curated instances across 14 fine-grained task types, covering semantic discrimination,…
arxiv.orgHowever, conventional function-calling benchmarks mainly evaluate task completion and API correctness, while privacy evaluation benchmarks typically focus on final responses or privacy…
arxiv.orgTUA-BenchUncategorized
We introduce TUA-Bench, a general-purpose benchmark for terminal-use agents.
arxiv.orgTo address this gap, we introduce DiscoBench, a benchmark for clarification-aware deep search, designed to evaluate whether search agents can proactively identify ambiguity, ask effective…
arxiv.orgTo study this, we introduce AIMS, a human-annotated dataset of 1,724 difficult safety prompts, each paired with an intent description and harm label.
arxiv.orgData from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanced accounts of need.
arxiv.orgCareQA-VisionUncategorized
Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinical and biomedical applications.
arxiv.orgWe introduce DMV-Bench (Code: https://github.
arxiv.orgTo address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit…
arxiv.orgWe introduce GAVEL (Grounded Caption Error Verification and Localization), a task that jointly addresses verification, explanation, and localization for image-text pairs.
arxiv.orgCurrent frameworks measure exclusively whether a model flags a video correctly rather than explaining why, turning evaluation into a black box where models can succeed through superficial…
arxiv.orgKo-WideSearchUncategorized
Web-agent benchmarks overwhelmingly measure depth -- pinning one obscure answer behind a chain of constraints -- while breadth, exhaustively enumerating a closed set and filling each…
arxiv.orgExisting red-teaming approaches are typically surface-specific and often recycle known attack templates; on text-poisoning benchmarks we measure 73-84% exact duplication.
arxiv.orgObviousBenchIntelligence & Reasoning
ObviousBench measures whether language models avoid simple, high-visibility mistakes — the kind of "obvious" errors users notice immediately.
obviousbench.comOmniRef-BenchMultimodal
To better assess model performance on complex MRIG tasks, we introduce OmniRef-Bench, a benchmark that covers complex combinations of reference image types and a large number of reference…
arxiv.orgTo address this gap, we introduce RedVox, a multilingual safety and fairness benchmark for audio and speech built on real voices, covering unsafe and unfair stereotypical requests across…
arxiv.orgWe introduce the Relational Stress and Psychiatry Corpus (RSPC) containing 1,799 Reddit posts annotated by psychiatrists for diagnostic categories, including the most prevalent mood…
arxiv.orgExisting AI-biology benchmarks largely measure broad knowledge, executable workflows, or local analysis steps.
arxiv.orgWe introduce SocialPersona, a benchmark for evaluating whether multimodal large language models (MLLMs) can recover revealed preferences from longitudinal social-media timelines and use…
arxiv.orgIn this paper, we introduce LCS-Bench, a stand-alone, theory-scale benchmark based on Logics for Computer Science.
arxiv.orgIn this paper, we take a step toward a broader evaluation by shifting the observable axis to execution resources: alongside test outcome and exception class, we predict peak memory,…
arxiv.orgWe introduce and formalize the problem of agentic surveillance: the ability of an AI agent to analyze available information, craft a report, and send it out using available tools.
arxiv.orgWe introduce ToolBench-X, a benchmark for evaluating agents under recoverable reliability hazards.
arxiv.orgTo fill this gap, we propose C3-Bench, a comprehensive benchmark for evaluating Context-aware Change Captioning.
arxiv.orgCyberChainBenchCybersecurity
We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability detection, exploit generation, and patch…
arxiv.orgTo more accurately characterize these capabilities, we propose a two-stage decoupled evaluation framework that separates exploit execution from reconnaissance.
arxiv.orgWe introduce DualEval, a latent model-item calibration framework that represents models and evaluation items in a shared space, jointly estimating model ability together with item…
arxiv.orgFacet-ProbeMultimodal
We introduce Facet-Probe, a five-facet audit (option, evidence-chunk, document-rank, image-set, and mixed-modality ordering) of 18 frontier and open-weight MLLMs.
arxiv.orgHowever, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in…
arxiv.orgWe introduce the Generalization Spectrum, an evaluation framework designed to expose this hidden dimension.
arxiv.orgHieroglyphBench evaluates vision-language models on transcribing photographed ancient Egyptian inscriptions into ordered Gardiner sign-list codes.
boggs.techWe present an empirical study of tool-augmented LLM agents on real-world energy market analytics tasks.
arxiv.orgLarge language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately reconstruct and apply the specific procedural decision…
arxiv.orgTo systematically evaluate this phenomenon, we introduce LibEvoBench, a multi-task benchmark spanning multiple versions of widely used Python libraries, along with a new metric, the…
arxiv.orgTo systematically study this problem, we introduce OCR-Robust, a benchmark designed for evaluating OCR reasoning robustness under visual perturbations.
arxiv.orgPQSG evaluates generated videos by checking their faithfulness to a prompt across objects, actions, and adherence to physical laws using a graph-based hierarchy of questions generated by a…
arxiv.orgHere we introduce Sarashina2.
arxiv.orgWe introduce \textsc{SpeechEQ}, a comprehensive framework designed to evaluate the sociolinguistic reasoning of Speech-Language Models (SLMs).
arxiv.orgSTEBTranslation
We introduce STEB (Speech-to-Speech Translation Expressiveness Benchmark), a 32.
arxiv.orgSWE-ProUncategorized
Unlike previous benchmarks, SWE-Pro pairs each task with parameterized tests to evaluate runtime, peak memory, and Time-Weighted Memory Usage (TWMU) across varying input data and execution…
arxiv.orgWe introduce TriViewBench, a controlled three-view visual reasoning benchmark constructed from synthetic 3D scenes with explicitly parameterized object count and occlusion.
arxiv.orgWe present Argus, a cross-regime benchmark for post-hoc UQ in single-step executable GUI grounding: a 27-method open-weight matrix over 4 VLM agents and 4 datasets, plus an 8-method…
arxiv.orgVideo carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single current image,…
arxiv.orgWBCMor VQAUncategorized
To address this limitation, we introduce WBCMor VQA, a clinically validated bilingual English, Urdu morphology aware VQA benchmark for leukemia and normal white blood cell analysis.
arxiv.orgWe introduce Age of LLM, a turn-based 1v1 benchmark in which two LLMs face off on a 13x7 grid to destroy the enemy base.
arxiv.orgAgentWorldBench evaluates how faithfully language world models predict the next environment observation after an agent action.
qwen.aiWe introduce Agora, a benchmark pairing 362 questions with eight domain collections of 9,664 authentic documents and 372M tokens, far exceeding any model's context window, so agents must…
arxiv.orgWe introduce BCoughBench, evaluating five FMs (OPERA-CT/CE/GT, HeAR, M2D+Resp) on nine classification tasks (AUROC, sensitivity at 95% specificity, Expected Calibration Error) and three…
arxiv.orgWe introduce BehaviorBench, a comprehensive benchmark that evaluates foundation models along four core capabilities: (1) behavior prediction and simulation, (2) strategic decision-making,…
arxiv.orgWe introduce AgenticInterpBench, a benchmark for circuit explanation built from 84 semi-synthetic transformer circuits with 163 component-level annotations.
arxiv.orgWe propose DramaDirector, a geometry-grounded framework that lets the planner borrow cinematographic geometry from a gallery of real short-drama shots indexed by depth and pose.
arxiv.orgWe introduce a dataset of 32,534 double-marked real student responses to GCSE mock exams (GCSEs are the UK's national exams, taken at age ~16), spanning 328 questions across five subjects…
arxiv.orgWe evaluate…
arxiv.orgMedBench v5Biomedicine
We introduce MedBench v5, a redesigned benchmark for clinical multimodal models (language, vision-language, and agent systems) that moves from static QA to dynamic, process-oriented…
arxiv.orgBuilt on synthetic ground truth for efficient, scalable measurement, MEMPROBE spans 50 simulated users with 31 hidden dimensions each (1,550 recovery targets) and tests 5 representative…
arxiv.orgHowever, existing benchmarks evaluate these only in isolation, leaving the interaction between biomedical expertise and multilingual coverage unmeasured.
arxiv.orgWe introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move beyond…
arxiv.orgWe introduce ParaPairAudioBench, a pairwise benchmark of 5,175 audio pairs across five paralinguistic dimensions: Style, Rate, Emphasis, Age, and Gender.
arxiv.orgWe present ReMMD, a realistic multilingual multi-image agentic verification framework for multimodal misinformation detection.
arxiv.orgWe present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity.
arxiv.orgMotivated by this finding, we propose Agent-as-a-Router, a framework that formalizes routing as a C-A-F loop (Context->Action->Feedback->Context).
arxiv.orgAgentCIBenchAgentic
Hence, we introduce AgentCIBench, an evaluation harness that turns this risk into executable, deterministically scored scenarios.
arxiv.orgEHR-ComplexBiomedicine
In this work, we introduce EHR-Complex, a large-scale benchmark designed for interactive clinical database reasoning.
arxiv.orgWe introduce EnterpriseClawBench, an enterprise agent benchmark constructed from proprietary, real-world agent sessions.
arxiv.orgWe introduce a matched execution-layer benchmark of 440 desktop tasks across 18 applications and 12 workflow categories, where screen-only GUI agents and skill-mediated CLI agents receive…
arxiv.orgHAKARI-BenchRetrieval
We present HAKARI-Bench, a lightweight benchmark that reconstructs existing retrieval suites into small datasets (Nano-sets): 35 benchmarks and 551 tasks across 43 languages in a unified…
arxiv.orgWe introduce HOLMES (Higher-Order Logic Meets real-world Explainable Symbolic reasoning), the first real-world benchmark for higher-order symbolic reasoning in LLMs, containing 1379…
arxiv.orgMotivated by this challenge, we introduce root memory, a structured, decision-preserving representation that distills reusable personalized logic from long-term user histories.
arxiv.orgThat is why we introduce IPO Finance Agent, which extends the Finance Agent v2 framework along two directions: task domain and retrieval architecture.
arxiv.orgWe introduce AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks, roles,…
arxiv.orgMulti-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based agents, each assigned a system prompt and a position within a workflow that governs…
arxiv.orgTo measure this phenomenon, we introduce TF-RefusalBench, a multilingual benchmark for criminal-law translation and summarization derived from public Swiss Supreme Court rulings.
arxiv.orgTo systematically investigate these hallucinations, we introduce MotionHalluc, a dedicated benchmark for evaluating motion hallucinations in paired-video comparison.
arxiv.orgWe introduce MuPPET (Multi-Party Privacy Exposure Testing), a benchmark for contextual privacy in multi-party conversations.
arxiv.orgTo address this gap, we introduce PIVOTS, the first benchmark built from Social-IQ 2.
arxiv.orgTo address this gap, we introduce RIFT-Bench, a graph representation-driven methodology for dynamic red-teaming that enables unified evaluations across diverse agentic architectures.
arxiv.orgSingGuard-BenchUncategorized
We present \textbf{SingGuard}, a policy-adaptive multimodal guardrail model family for safety assessment in multimodal conversations.
arxiv.org
New ModelsAll
August 2026
MiniMax H3MiniMax
Video Generation · Open weights
huggingface.coReasoning model · Open weights announced · $2 in / $6 out per 1M tokens
qwen.ai
July 2026
Reasoning model · Open weights · $0.14 in / $0.28 out per 1M tokens
api-docs.deepseek.comINInkling-SmallThinking Machines Lab
Reasoning model · Open weights
thinkingmachines.aiClaude Opus 5Anthropic
Reasoning model · API-only · $5 in / $25 out per 1M tokens
anthropic.comReasoning model · API-only · $1.5 in / $7.5 out per 1M tokens
deepmind.googlePOLaguna S 2.1Poolside
Language model · API-only
poolside.aiImage Generation · API-only
qwen.aiReasoning model · API-only
qwen.aiKIKimi K3Moonshot AI
Reasoning model · Open weights · $3 in / $15 out per 1M tokens
kimi.comINInklingThinking Machines Lab
Reasoning model · Open weights · $1.87 in / $4.68 out per 1M tokens
thinkingmachines.aiReasoning model · API-only · $1 in / $6 out per 1M tokens
openai.comGPT-5.6 SolOpenAI
Reasoning model · API-only · $5 in / $30 out per 1M tokens
openai.comReasoning model · API-only · $2.5 in / $15 out per 1M tokens
openai.comReasoning model · API-only
ai.meta.comReasoning model · API-only · $2 in / $6 out per 1M tokens
x.aiReasoning model · API-only · $2.5 in / $12.5 out per 1M tokens
cognition.comImage Generation · Open weights
research.nvidia.comReasoning model · Open weights
huggingface.co
June 2026
Claude Sonnet 5Anthropic
Reasoning model · API-only · $3 in / $15 out per 1M tokens
anthropic.comClosed/API · $0.25 in / $1.5 out per 1M tokens
openrouter.aiClosed/API · $5 in / $30 out per 1M tokens
openrouter.aiReasoning model · Open weights
huggingface.coClosed/API · $0.5 in / $3 out per 1M tokens
openrouter.aiClosed/API · $2 in / $12 out per 1M tokens
openrouter.aiClosed/API · $0 in / $0 out per 1M tokens
openrouter.aiReasoning model · Open weights · $0.94 in / $3 out per 1M tokens
z.aiFusionOpenRouter
Closed/API
openrouter.aiKIKimi K2.7 CodeMoonshot AI
Reasoning model · Open weights · $0.74 in / $3.5 out per 1M tokens
huggingface.coClaude Fable 5Anthropic
Reasoning model · API-only · $10 in / $50 out per 1M tokens
anthropic.comClosed/API · $10 in / $50 out per 1M tokens
openrouter.aiReasoning model · API-only
anthropic.comNANex-N2-MiniNex-AGI
Reasoning model · Open source
nex-agi.comReasoning model · Open weights · $0.25 in / $1 out per 1M tokens
nex-agi.comReasoning model · API-only
cognition.aiGemma 4 E2BGoogle
Reasoning model · Open weights, gated
huggingface.coGemma 4 E4BGoogle
Reasoning model · Open weights, gated
huggingface.coLanguage model · Open source
magenta.withgoogle.comReasoning model · Open weights · $0.5 in / $2.2 out per 1M tokens
huggingface.coClosed/API · $0 in / $0 out per 1M tokens
openrouter.aiGemma 4 12BGoogle
Language model · Open weights, gated
blog.googleReasoning model · API-only · $0.32 in / $1.28 out per 1M tokens
openrouter.aiLanguage model · API-only
microsoft.aiMAI-Image-2.5Microsoft
Language model · API-only · $5 in / $47 out per 1M tokens
microsoft.aiLanguage model · API-only · $1.75 in / $19.5 out per 1M tokens
microsoft.aiMAI-Thinking-1Microsoft
Reasoning model · API-only
microsoft.aiLanguage model · API-only
microsoft.aiMAI-Voice-2Microsoft
Language model · API-only
microsoft.ai