Overall
deepseek-v4p1-flash
fireworks · best observeddeepseek-v4p1-flash has the highest observed rating. 3 selected evidence items.
bench_model_ratings · 21 Sep 2026, 20:10 UTCAPTEVA BENCHMARKS
Model and provider results across agent workloads, with dated benchmark runs and a rolling view of performance over time.
01 / BENCHMARK OVERVIEW
Category-balanced overall rating and the underlying Coding, Marketing, and Analytics ratings.
Overall
deepseek-v4p1-flash has the highest observed rating. 3 selected evidence items.
bench_model_ratings · 21 Sep 2026, 20:10 UTCCoding
deepseek-v4p1-flash has the highest observed rating. 1 selected evidence item.
run_1790019834873287670 · 21 Sep 2026, 20:10 UTCMarketing
gpt-5.6-luna has the highest observed rating. 1 selected evidence item.
run_1790019834863080849 · 21 Sep 2026, 19:51 UTCAnalytics
gpt-5.6-sol has the highest observed rating. 1 selected evidence item.
run_1790019834865991645 · 21 Sep 2026, 19:57 UTC“Best” means the highest observed score under Category-balanced model rating v3. Overall gives each category equal weight. Coding and Analytics use 80% deterministic correctness plus 20% efficiency; Marketing retains the source quality score. These small samples do not establish a universal winner.
02 / BENCHMARK MOVEMENT
Source benchmark scores from successive completed runs in the selected category. A directional signal, not a definitive ranking.
deepseek-v4p1-flash
1 comparable window available
gpt-5.6-sol
2 comparable windows available
glm-5p3
1 comparable window available
gpt-5.6-terra
1 comparable window available
gpt-5.6-luna
2 comparable windows available
kimi-k3
1 comparable window available
CURRENT OBSERVED CHOICE / CODING
deepseek-v4p1-flash has the highest current category rating at 87.5. The sample is small, so this is a practical routing preference rather than a proven universal lead.
Published snapshot · Bench installation 189 · 049871636cd9Movement shows returned source scores for the current sealed pack. Replacement evidence can change a model’s current rating without rewriting earlier runs. Data becomes stale here after 30 minutes; that threshold is a page policy, not a provider incident claim.
03 / DETAILED RESULTS
The published rating snapshot, its selected evidence, completion time, and source task outcome.
| Provider / model | Rating 0–100 · policy score | Source outcome passed · selected evidence | Average completion seconds | Average tokens per selected result | Evidence |
|---|---|---|---|---|---|
| deepseek-v4p1-flashfireworks | 80.1Equal mean across categories | 1 / 333.3% source task pass rate | 96.3 s | 280,434 | Selected3 evidence items · 0 invalid |
| gpt-5.6-solopenai-codex | 76.8Equal mean across categories | 1 / 333.3% source task pass rate | 220.5 s | 275,641 | Selected3 evidence items · 0 invalid |
| glm-5p3fireworks | 76.0Equal mean across categories | 1 / 333.3% source task pass rate | 318.2 s | 453,386 | Selected3 evidence items · 0 invalid |
| gpt-5.6-terraopenai-codex | 75.4Equal mean across categories | 1 / 333.3% source task pass rate | 248.8 s | 265,708 | Selected3 evidence items · 0 invalid |
| gpt-5.6-lunaopenai-codex | 74.8Equal mean across categories | 1 / 333.3% source task pass rate | 213.0 s | 268,949 | Selected3 evidence items · 0 invalid |
| kimi-k3fireworks | 69.0Equal mean across categories | 0 / 20.0% source task pass rate | 187.7 s | 219,130 | Incomplete2 evidence items · 1 invalid |
No universal winner is claimed. These ratings use the latest admitted result per model, pack, scenario, and trial. 20 source results remain preserved; 17 latest comparable results contribute to the displayed rating.
04 / THE METHOD BEHIND THE NUMBERS
Ratings remain tied to sealed benchmark packs, an immutable rating policy, dated runs and explicit evidence-selection rules.
Use sealed pack and scoring-profile digests. Record the exact model, provider, environment, scenario definitions, budgets and run timestamps.
Run versioned scenarios with deterministic checks and, where configured, a pinned judge model and rubric. Keep each provider in the same category and pack.
Use the latest admitted result per model, pack, scenario and trial. Preserve superseded source results so the displayed rating remains auditable.
Show selected and preserved result counts, source task outcomes, average completion time, token use, missing evidence and policy caveats beside every rating.
The Web SDK calls bench_model_ratings, bench_category_list, bench_pack_list and bench_run_search. The rating tool supplies the category-balanced overall score; completed runs supply timestamps, source outcomes, duration, tokens and movement.
Current source: Published production snapshot from Bench v0.10.0 installation 189. Live source checks remain available.
Each bar is a returned source benchmark score for the same sealed pack and model, ordered by completion time. Replacement evidence is retained as another window. The page does not interpolate missing windows or translate a failed task into a provider outage.
Category-balanced model rating v3 gives Coding, Marketing and Analytics equal weight. Coding and Analytics combine 80% deterministic correctness with 20% efficiency. Marketing uses the source Agentic Quality v2 score. The latest admitted result per model, pack, scenario and trial is selected.
The displayed rating must have complete selected evidence. The highest rating is the current observed choice; exact ties remain ties. Small samples and stale data prevent a definitive universal ranking.
This view uses aggregate results and source identifiers. A public deployment should use a read-only scoped credential or a server-side redaction boundary. Prompts, directives, private environment identifiers and full evidence bundles should remain scoped.
APTEVA
Run persistent agents with the models, tools and deployment setup that fit your workload.
7 days free, no credit card. Or self-host the open-source platform.