Marble AI agent benchmark

Glass Bio

4 min read Original article ↗

Improving the performance of low-cost open models using tool calls

We believe that the right tools, the right harness, and access to compute will allow open-source models to reach parity with frontier models in their ability to help scientists. This series of blogposts will document our progress towards finding the best configuration for this task.

We start by building a set of evaluations that test bio reasoning ability. Our suite contains 50 paper follow-up cases: one paper builds or rejects findings of a prior paper. Can a model interpret an original paper’s anonymized evidence, anticipate the concerns later raised by a follow-up paper, and propose experiments that could resolve them?

We run our evals through two models: GPT 5.6 Sol and DeepSeek v4 Flash. We also compare runs with and without access to web search or the FireCrawl Research Index API. Example questions are shown in the appendix.

Paper 1

Follow-up Paper

Blinded Paper 1 data

Model predicts follow-up findings

True follow-up findings

Compare and score

Model performance by cost

GPT 5.6 Sol outscores DeepSeek V4 Flash, both with and without search. Across search conditions, blinding lowers DeepSeek's average by 0.61 points and GPT's by 0.52 points. GPT is also more consistent across repeated runs than DeepSeek. Across all GPT conditions, 4.9% of attempts were blocked by guardrails, despite questions stemming from published work.

Average rubric score across 25 cases against cost per 10 questions for benchmark runs.5.05.56.06.57.07.58.08.59.09.5$0.10$0.20$0.50$1$2$5$10$20Cost per 10 questionsAverage rubric score across 25 casesDeepSeekDeepSeek +Firecrawl searchDeepSeek +web searchGPT 5.6 SolGPT 5.6 Sol +web searchGPT 5.6 Sol +Firecrawl search

Non-blinded ×Blinded

Firecrawl closes part of the DeepSeek to GPT score gapDeepSeek scores 5.71 without search and 6.33 with Firecrawl. GPT 5.6 Sol scores 8.19 without search. Firecrawl closes 25 percent of the original gap.DeepSeek5.71DeepSeek + Firecrawl6.33GPT 5.6 Sol8.19+0.62 points1.86-point gap

What does Firecrawl add?

Firecrawl raises DeepSeek's average from 5.71 to 6.33, a paired gain of 0.62 points. This reduces the gap to GPT 5.6 Sol by 25% and costs $0.72 per 10 questions, compared with $3.63 for GPT.

Can LLMs identify paper of origin based on data alone?

We also ask the LLM whether it can, based on the data, recognize the paper of origin. Anonymizing the data by removing identifiers (genes, molecules, tissues, and more) sharply reduces recognition for DeepSeek, but much less so for GPT 5.6 Sol. In blinded runs, the Firecrawl API raises DeepSeek's recognition rate from 29% to 51%. Paper recognition rate also strongly correlates with performance in non-blinded runs (r = 0.80) and blinded runs (r = 0.94). This could be attributed both to a gradient in prompt difficulty, or partial memorization of literature rather than case-by-case reasoning.

Average rubric score against paper-recognition rate for 36 benchmark runs. Separate lines show the correlation for blinded and non-blinded runs.5.56.06.57.07.58.08.59.09.520%30%40%50%60%70%80%90%100%Paper-recognition rateAverage rubric score across 25 casesDeepSeekDeepSeek +Firecrawl searchDeepSeek +web searchGPT 5.6 SolGPT 5.6 Sol +Firecrawl searchGPT 5.6 Sol +web search

Non-blinded ×Blinded

When does Firecrawl improve scientific scores?

DeepSeek's median score increases from 5 to 7 when it recognizes the source paper, while GPT 5.6 Sol's median increases from 7 to 10. In matched DeepSeek runs, Firecrawl adds 0.62 points overall. The gain is 0.27 when neither condition recognizes the paper. It reaches 0.83 when only Firecrawl recognizes it and 0.78 when both conditions recognize it. The 0.27-point gain is modest and not statistically significant. The 0.78-point gain suggests that metadata and literature text provide useful evidence beyond paper identification.

Firecrawl score gain by paper-recognition outcomeFirecrawl adds 0.62 points across all cases. The paired gain is 0.27 when neither condition recognizes the paper, 0.83 when only Firecrawl recognizes it, 0.78 when both conditions recognize it, and 0.13 when only the no-search condition recognizes it.-0.50+0.00+0.50+1.00+1.50All casesn=148+0.62Never recognizedn=41+0.27Recognized only with Firecrawln=36+0.83Recognized with and without Firecrawln=63+0.78Not recognized only with Firecrawln=8+0.13Paired score gain with Firecrawl

Overall, adding search capabilities improves benchmark performance while keeping costs low. In the coming weeks, we'll release more experiments as we improve our internal stack and push open models closer to frontier models.

We're looking for test users! hello@glass.bio if you're interested in trying out our tools.

Appendix: example questions and responses

Challenge

I am studying human Ins(1,3,4)P3 5/6-kinase. My preparations show the expected inositol kinase activity and also phosphorylate ATF-2. I just received the fractionation, immunodepletion, and recombinant-protein results in study_1_figure_5.jpg and study_1_figure_6.jpg. Their complete source captions are in study_1_figure_captions.txt, and the matching source Results section is in study_1_results.txt.

Can you analyze these results and tell me what you think is going on? If you recognize the source paper, please identify it.