APTEVA BENCHMARKS

Benchmarks

Model and provider results across agent workloads, with dated benchmark runs and a rolling view of performance over time.

Benchmarks
current & historical

01 / BENCHMARK OVERVIEW

Model ratings by benchmark type

Category-balanced overall rating and the underlying Coding, Marketing, and Analytics ratings.

Latest published results
3 benchmark categories20 preserved source results17 latest comparable results selectedIncomplete evidence excluded from “best observed”

Overall

deepseek-v4p1-flash

fireworks · best observed
80.1

deepseek-v4p1-flash has the highest observed rating. 3 selected evidence items.

bench_model_ratings · 21 Sep 2026, 20:10 UTC

Coding

deepseek-v4p1-flash

fireworks · best observed
87.5

deepseek-v4p1-flash has the highest observed rating. 1 selected evidence item.

run_1790019834873287670 · 21 Sep 2026, 20:10 UTC

Marketing

gpt-5.6-luna

openai-codex · best observed
96.5

gpt-5.6-luna has the highest observed rating. 1 selected evidence item.

run_1790019834863080849 · 21 Sep 2026, 19:51 UTC

Analytics

gpt-5.6-sol

openai-codex · best observed
57.6

gpt-5.6-sol has the highest observed rating. 1 selected evidence item.

run_1790019834865991645 · 21 Sep 2026, 19:57 UTC

“Best” means the highest observed score under Category-balanced model rating v3. Overall gives each category equal weight. Coding and Analytics use 80% deterministic correctness plus 20% efficiency; Marketing retains the source quality score. These small samples do not establish a universal winner.

02 / BENCHMARK MOVEMENT

Recent benchmark movement

Source benchmark scores from successive completed runs in the selected category. A directional signal, not a definitive ranking.

Category
FreshApteva Code · Production issue repair v2 v1.0.0Observed 21 Sep 2026, 20:10 UTC
Latest selected evidence run_1790019834873287670. Each bar is one returned source-score window; missing windows are omitted.
FStable

fireworks

deepseek-v4p1-flash

Current rating
87.5
Average completion
120.2 s
Source outcome
0 / 1 passed

1 comparable window available

OImproving

openai-codex

gpt-5.6-sol

Current rating
79.3
Average completion
388.6 s
Source outcome
0 / 1 passed

2 comparable windows available

FStable

fireworks

glm-5p3

Current rating
76.3
Average completion
555.8 s
Source outcome
0 / 1 passed

1 comparable window available

OStable

openai-codex

gpt-5.6-terra

Current rating
72.3
Average completion
429.0 s
Source outcome
0 / 1 passed

1 comparable window available

OImproving

openai-codex

gpt-5.6-luna

Current rating
71.2
Average completion
393.6 s
Source outcome
0 / 1 passed

2 comparable windows available

FStable

fireworks

kimi-k3

Current rating
82.0
Average completion
252.5 s
Source outcome
0 / 1 passed

1 comparable window available

CURRENT OBSERVED CHOICE / CODING

deepseek-v4p1-flash for this measured category. kimi-k3 is the next observed alternative.

deepseek-v4p1-flash has the highest current category rating at 87.5. The sample is small, so this is a practical routing preference rather than a proven universal lead.

Published snapshot · Bench installation 189 · 049871636cd9

Movement shows returned source scores for the current sealed pack. Replacement evidence can change a model’s current rating without rewriting earlier runs. Data becomes stale here after 30 minutes; that threshold is a page policy, not a provider incident claim.

03 / DETAILED RESULTS

Rating evidence

The published rating snapshot, its selected evidence, completion time, and source task outcome.

Dated snapshot
bench_model_ratings · 21 Sep 2026, 20:10 UTC17 selected of 20 preserved source results
Published overall rating · dated snapshot
Provider / modelRating 0–100 · policy scoreSource outcome passed · selected evidenceAverage completion secondsAverage tokens per selected resultEvidence
deepseek-v4p1-flashfireworks80.1
Equal mean across categories
1 / 333.3% source task pass rate96.3 s280,434Selected3 evidence items · 0 invalid
gpt-5.6-solopenai-codex76.8
Equal mean across categories
1 / 333.3% source task pass rate220.5 s275,641Selected3 evidence items · 0 invalid
glm-5p3fireworks76.0
Equal mean across categories
1 / 333.3% source task pass rate318.2 s453,386Selected3 evidence items · 0 invalid
gpt-5.6-terraopenai-codex75.4
Equal mean across categories
1 / 333.3% source task pass rate248.8 s265,708Selected3 evidence items · 0 invalid
gpt-5.6-lunaopenai-codex74.8
Equal mean across categories
1 / 333.3% source task pass rate213.0 s268,949Selected3 evidence items · 0 invalid
kimi-k3fireworks69.0
Equal mean across categories
0 / 20.0% source task pass rate187.7 s219,130Incomplete2 evidence items · 1 invalid

No universal winner is claimed. These ratings use the latest admitted result per model, pack, scenario, and trial. 20 source results remain preserved; 17 latest comparable results contribute to the displayed rating.

04 / THE METHOD BEHIND THE NUMBERS

Make every result inspectable.

Ratings remain tied to sealed benchmark packs, an immutable rating policy, dated runs and explicit evidence-selection rules.

01

Freeze the experiment

Use sealed pack and scoring-profile digests. Record the exact model, provider, environment, scenario definitions, budgets and run timestamps.

02

Measure actual work

Run versioned scenarios with deterministic checks and, where configured, a pinned judge model and rubric. Keep each provider in the same category and pack.

03

Select evidence explicitly

Use the latest admitted result per model, pack, scenario and trial. Preserve superseded source results so the displayed rating remains auditable.

04

Keep the sample visible

Show selected and preserved result counts, source task outcomes, average completion time, token use, missing evidence and policy caveats beside every rating.

What does this page load?

The Web SDK calls bench_model_ratings, bench_category_list, bench_pack_list and bench_run_search. The rating tool supplies the category-balanced overall score; completed runs supply timestamps, source outcomes, duration, tokens and movement.

Current source: Published production snapshot from Bench v0.10.0 installation 189. Live source checks remain available.

How is recent movement calculated?

Each bar is a returned source benchmark score for the same sealed pack and model, ordered by completion time. Replacement evidence is retained as another window. The page does not interpolate missing windows or translate a failed task into a provider outage.

How is the overall rating calculated?

Category-balanced model rating v3 gives Coding, Marketing and Analytics equal weight. Coding and Analytics combine 80% deterministic correctness with 20% efficiency. Marketing uses the source Agentic Quality v2 score. The latest admitted result per model, pack, scenario and trial is selected.

When is a recommendation allowed?

The displayed rating must have complete selected evidence. The highest rating is the current observed choice; exact ties remain ties. Small samples and stale data prevent a definitive universal ranking.

What is public, and what stays scoped?

This view uses aggregate results and source identifiers. A public deployment should use a read-only scoped credential or a server-side redaction boundary. Prompts, directives, private environment identifiers and full evidence bundles should remain scoped.

APTEVA

Put the right model to work.

Run persistent agents with the models, tools and deployment setup that fit your workload.

7 days free, no credit card. Or self-host the open-source platform.