Intelligence and Inference Performance Summary
Average Score (16K max context) vs. End-to-End Generation TimeiPhone 17 Pro
Simple average of 5 evaluations chosen to represent real-world mobile device usage: BFCL (subset), IFBench, AA-Omniscience, GPQA Diamond, MATH-500 · Context limited to 16K tokens · E2E time is the seconds taken to process a 1024-token prompt and generate a 256-token response
A simple average of five evaluations chosen to represent real-world mobile device usage, run against models small enough to fit on portable hardware and each measured independently by Artificial Analysis. See the methodology for further details.
Total wall-clock time to process a 1,024-token prompt and generate a 256-token response. Note that this is fundamental to the hardware used and the model’s architecture, and does not include the effect of model verbosity or tendency to use more or fewer turns.
Inference Performance
End-to-End Generation TimeiPhone 17 Pro
Seconds to process a 1024-token prompt and generate a 256-token response · Lower is better
Total wall-clock time to process a 1,024-token prompt and generate a 256-token response. Note that this is fundamental to the hardware used and the model’s architecture, and does not include the effect of model verbosity or tendency to use more or fewer turns.
Model Intelligence
Average Score (Mobile Device Benchmark Set, 16K max context)iPhone 17 Pro
Simple average of 5 evaluations chosen to represent real-world mobile device usage: BFCL (subset), IFBench, AA-Omniscience, GPQA Diamond, MATH-500 · Context limited to 16K tokens · Higher is better
Performance measurements omitted as the model did not fit on this device or exceeded the time limit
A simple average of five evaluations chosen to represent real-world mobile device usage, run against models small enough to fit on portable hardware and each measured independently by Artificial Analysis. See the methodology for further details.
Token Efficiency
Context Budget Overruns
Generations that stopped at the 16K-token limit instead of finishing, across every evaluation run on the model · Lower is better
Evaluation Breakdown
Mobile Device Benchmark Set Evaluations (16K max context)iPhone 17 Pro
Intelligence evaluations measured independently by Artificial Analysis · Context limited to 16K tokens · Higher is better
BFCL
Tool calling (index subset)
IFBench
Instruction following
AA-Omniscience Accuracy
Knowledge
AA-Omniscience Non-Hallucination Rate
1 - hallucination rate
GPQA Diamond
Scientific reasoning
MATH-500
Quantitative reasoning