AgenticVBench - Video Post-Production Benchmark for AI Agents

2 min read Original article ↗

May 2026

Can AI agents do real-world video post-production work?

We gave the 10 best frontier models 100 expert-authored tasks across the four stages of video post-production. The best agent tops out at 38%. Human experts scored 89%.

Why this benchmark exists

Verification is not here for free.

RLVR - reinforcement learning with verifiable rewards - works in math and code because centuries of humanistic work built the verifiers; the bill was paid before we got there. Creative work hasn't paid that bill. AgenticVBench is what paying it looks like in film.

It also measures the sim2real gap: the distance between how agents score on tidy lab benchmarks and how they hold up on real post-production work. Here that gap is stark - the best frontier agent scores 38%, human experts 89%.

Read the full essay →

What the bench tests

Four task families spanning the real-world post-production workflow.

Authored by 20 industry experts averaging 6 years of post-production experience. Tasks span 30 minutes to one week of human work.

The harness finding

The harness matters as much as the model.

Holding the model fixed and varying the harness shifts GPT-5.5's Assembly score by 20 percentage points, comparable to the gap between adjacent models on the leaderboard.

Most benchmarks today are still model-based. The data here says that's wrong. Agent performance is determined by both the model and the scaffolding around it. Reporting only the model misses the larger story.

Agent = model × harness.

GPT-5.5 on Assembly · score by harness

Same model. 20-point swing.

Cite this work

Citation

If you find AgenticVBench useful in your research, please consider citing the paper.

BibTeX

@article{cao2026agenticvbench,
  title={AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?},
  author={Cao, Zongheng and Zheng, Yi and Song, Rui and Hu, Xinyu},
  journal={arXiv preprint arXiv:2605.27705},
  year={2026}
}

Read the paper on arXiv →