Chan Lab

REAL DATA / HISTORICAL EXAMPLE / 21 SEP 2026

What a real run looks like.

A few records. The full context. Nothing made up.

Markdown + Mermaid ↓Summary & sample JSON ↓20 selected records · CSV ↓

What this example shows

Two real experiments from 21 September 2026: 50 calculus questions and 50 decision scenarios, each with three option orders and four API/input configurations. Each experiment contains 600 attempts. No new model calls were made to create this example.

Historical data, not your live score

These measurements came from the earlier local Chrome/Playwright experiment, not the current browser-only runner. The question renderer, network path, hardware and execution protocol differ. Do not combine these numbers with a new Chan Lab run as if they were the same test.

How the excerpt was chosen

All summary values are recomputed from the complete 1,200 attempts and checked against the original verification files. Only 20 attempt records are included here: the first shared question for all four APIs in each experiment, the valid request closest to each API’s median latency, and all four failed requests. This is a deterministic illustration, not a random sample; do not recompute overall scores from these 20 rows.

Setup and timing

Jev responses identified jev-1.13.0. OpenJev used AlexWortega/openjev/qwen3.5-4b-nli-v2 on an Apple M4 Pro with 48 GB unified memory, MPS and BF16. Text, screenshot and combined inputs were compared. The cloud and local deployments used different hardware and network paths. HTTP duration includes network and service processing through response parsing and validation; it is not pure model thinking time. P50/P95 use linear interpolation and include failed attempts.

Calculus — five sets

50 base questions × 3 rounds × 4 groups = 600 attempts. Each API’s fixed denominator is 150, including failures.

API / inputCorrect / 150AccuracyValid accuracyFailuresHTTP mean sP50 sP95 s
Jev text127/15084.7%84.7% (127/150)00.8630.6681.393
OpenJev text79/15052.7%52.7% (79/150)01.5971.6011.713
OpenJev image72/15048.0%48.0% (72/150)04.1634.2294.362
OpenJev text + image83/15055.3%55.3% (83/150)04.4124.4824.616
Calculus — five sets accuracyCalculus — five sets accuracyJev text84.67 %OpenJev text52.67 %OpenJev image48.00 %OpenJev text + image55.33 % Calculus — five sets mean API timeCalculus — five sets mean API timeJev text0.86 sOpenJev text1.60 sOpenJev image4.16 sOpenJev text + image4.41 s

Per-round data

API / inputRound 1 / 50Round 2 / 50Round 3 / 50
Jev text44/5040/5043/50
OpenJev text27/5024/5028/50
OpenJev image24/5024/5024/50
OpenJev text + image28/5026/5029/50

A question shared by all four APIs

S2-Q01-R1
For f(x) = 3x³ + 4x² − 5x, find f′(1).

A. 8
B. 2
C. 12
D. 17

Answer key: C

Selected attempt records

Global #API / inputQuestionChoice / goldResultHTTP secondsWhy included
1Jev textS2-Q01-R1D / Cincorrect3.077755first shared question
2OpenJev textS2-Q01-R1A / Cincorrect1.816309first shared question
3OpenJev imageS2-Q01-R1A / Cincorrect4.252545first shared question
4OpenJev text + imageS2-Q01-R1A / Cincorrect4.613546first shared question
138OpenJev text + imageS1-Q04-R1D / Dcorrect4.484471nearest valid latency median
158Jev textS4-Q04-R1B / Bcorrect0.668401nearest valid latency median
172OpenJev textS5-Q05-R1B / Aincorrect1.602863nearest valid latency median
198OpenJev imageS5-Q07-R1C / Ccorrect4.228928nearest valid latency median

In this calculus bank, Jev text scored 127/150 = 84.7%; OpenJev text, image and combined input scored 52.7%, 48.0% and 55.3%. No requests failed. These observed scores describe this fixed question bank and deployment, not a general ranking.

Decision scenarios — five sets

50 base questions × 3 rounds × 4 groups = 600 attempts. Each API’s fixed denominator is 150, including failures.

API / inputCorrect / 150AccuracyValid accuracyFailuresHTTP mean sP50 sP95 s
Jev text122/15081.3%83.6% (122/146)41.1480.7252.194
OpenJev text93/15062.0%62.0% (93/150)02.3322.2042.888
OpenJev image90/15060.0%60.0% (90/150)04.8154.7804.963
OpenJev text + image105/15070.0%70.0% (105/150)05.8755.7356.422
Decision scenarios — five sets accuracyDecision scenarios — five sets accuracyJev text81.33 %OpenJev text62.00 %OpenJev image60.00 %OpenJev text + image70.00 % Decision scenarios — five sets mean API timeDecision scenarios — five sets mean API timeJev text1.15 sOpenJev text2.33 sOpenJev image4.82 sOpenJev text + image5.87 s

Per-round data

API / inputRound 1 / 50Round 2 / 50Round 3 / 50
Jev text39/5043/5040/50
OpenJev text32/5030/5031/50
OpenJev image30/5032/5028/50
OpenJev text + image37/5035/5033/50

A question shared by all four APIs

D1-Q08-R1
School poster. One person; hands-on steps cannot overlap or pause. Alone phases can overlap anything.
Design: 2m hands.
Print: 1m hands, then 5m alone; after Design fully finish.
Frame: 4m hands.
Mount: 2m hands; after Print + Frame fully finish.
Tidy: 2m hands.
Follow each listed order, waiting if needed. Which finishes all work earliest?

A. Design → Tidy → Print → Frame → Mount
B. Design → Print → Frame → Mount → Tidy
C. Tidy → Design → Print → Frame → Mount
D. Design → Print → Tidy → Frame → Mount

Answer key: D

Selected attempt records

Global #API / inputQuestionChoice / goldResultHTTP secondsWhy included
1Jev textD1-Q08-R1B / Dincorrect0.937435first shared question
2OpenJev textD1-Q08-R1B / Dincorrect3.942378first shared question
3OpenJev imageD1-Q08-R1B / Dincorrect5.378662first shared question
4OpenJev text + imageD1-Q08-R1B / Dincorrect7.111950first shared question
78Jev textD2-Q03-R1— / Afetch failed [UND_ERR_CONNECT_TIMEOUT]10.380249all failed attempts
108OpenJev textD2-Q02-R1D / Dcorrect2.203423nearest valid latency median
150OpenJev imageD4-Q10-R1C / Dincorrect4.780264nearest valid latency median
152Jev textD4-Q10-R1D / Dcorrect0.716782nearest valid latency median
263OpenJev text + imageD5-Q09-R2B / Bcorrect5.736382nearest valid latency median
568Jev textD2-Q06-R3— / Bfetch failed [UND_ERR_CONNECT_TIMEOUT]10.361398all failed attempts
587Jev textD1-Q08-R3— / Bfetch failed [UND_ERR_CONNECT_TIMEOUT]10.368178all failed attempts
600Jev textD3-Q02-R3— / Dfetch failed [UND_ERR_CONNECT_TIMEOUT]10.385349all failed attempts

Four Jev calls ended in connection timeouts (UND_ERR_CONNECT_TIMEOUT), with no HTTP response. They count as unsuccessful attempts in 122/150 = 81.3%; valid-answer accuracy is 122/146 = 83.6%. These failures do not prove a reasoning error. No retries were substituted. OpenJev combined input scored 105/150 = 70.0%, with a mean HTTP duration of 5.875 seconds in this local deployment.

Interpretation and limits

The two banks are not pooled into one accuracy or timing score. Repeated option orders are not independent questions. There was no external difficulty calibration. The original evaluator clicked the returned option using Playwright; it did not test independent coordinate grounding. Cloud Jev and local OpenJev latency are not a controlled hardware comparison. The original screenshot sizes also differ from the current browser-rendered cards.

Publication and provenance

Only an explicit field allowlist is exported: API group, question IDs, selected/gold answers, correctness, HTTP status/error, UTC timestamps and stage durations. Credentials, Authorization headers, API endpoint URLs, raw request/response bodies, machine paths and internal document links are excluded. The historical measurements are static public examples; visitors’ new results remain in their browser.

The downloadable JSON preserves exact aggregate values and provenance hashes. The CSV preserves original UTC timestamps and all selected timing stages. Full-run source files are not included in this excerpt. Mermaid source for the four charts is in the Markdown download.

Try your own APIs →