REAL DATA / HISTORICAL EXAMPLE / 21 SEP 2026
What a real run looks like.
A few records. The full context. Nothing made up.
What this example shows
Two real experiments from 21 September 2026: 50 calculus questions and 50 decision scenarios, each with three option orders and four API/input configurations. Each experiment contains 600 attempts. No new model calls were made to create this example.
Historical data, not your live score
These measurements came from the earlier local Chrome/Playwright experiment, not the current browser-only runner. The question renderer, network path, hardware and execution protocol differ. Do not combine these numbers with a new Chan Lab run as if they were the same test.
How the excerpt was chosen
All summary values are recomputed from the complete 1,200 attempts and checked against the original verification files. Only 20 attempt records are included here: the first shared question for all four APIs in each experiment, the valid request closest to each API’s median latency, and all four failed requests. This is a deterministic illustration, not a random sample; do not recompute overall scores from these 20 rows.
Setup and timing
Jev responses identified jev-1.13.0. OpenJev used AlexWortega/openjev/qwen3.5-4b-nli-v2 on an Apple M4 Pro with 48 GB unified memory, MPS and BF16. Text, screenshot and combined inputs were compared. The cloud and local deployments used different hardware and network paths. HTTP duration includes network and service processing through response parsing and validation; it is not pure model thinking time. P50/P95 use linear interpolation and include failed attempts.
Calculus — five sets
50 base questions × 3 rounds × 4 groups = 600 attempts. Each API’s fixed denominator is 150, including failures.
| API / input | Correct / 150 | Accuracy | Valid accuracy | Failures | HTTP mean s | P50 s | P95 s |
|---|---|---|---|---|---|---|---|
| Jev text | 127/150 | 84.7% | 84.7% (127/150) | 0 | 0.863 | 0.668 | 1.393 |
| OpenJev text | 79/150 | 52.7% | 52.7% (79/150) | 0 | 1.597 | 1.601 | 1.713 |
| OpenJev image | 72/150 | 48.0% | 48.0% (72/150) | 0 | 4.163 | 4.229 | 4.362 |
| OpenJev text + image | 83/150 | 55.3% | 55.3% (83/150) | 0 | 4.412 | 4.482 | 4.616 |
Per-round data
| API / input | Round 1 / 50 | Round 2 / 50 | Round 3 / 50 |
|---|---|---|---|
| Jev text | 44/50 | 40/50 | 43/50 |
| OpenJev text | 27/50 | 24/50 | 28/50 |
| OpenJev image | 24/50 | 24/50 | 24/50 |
| OpenJev text + image | 28/50 | 26/50 | 29/50 |
A question shared by all four APIs
S2-Q01-R1 For f(x) = 3x³ + 4x² − 5x, find f′(1). A. 8 B. 2 C. 12 D. 17 Answer key: C
Selected attempt records
| Global # | API / input | Question | Choice / gold | Result | HTTP seconds | Why included |
|---|---|---|---|---|---|---|
| 1 | Jev text | S2-Q01-R1 | D / C | incorrect | 3.077755 | first shared question |
| 2 | OpenJev text | S2-Q01-R1 | A / C | incorrect | 1.816309 | first shared question |
| 3 | OpenJev image | S2-Q01-R1 | A / C | incorrect | 4.252545 | first shared question |
| 4 | OpenJev text + image | S2-Q01-R1 | A / C | incorrect | 4.613546 | first shared question |
| 138 | OpenJev text + image | S1-Q04-R1 | D / D | correct | 4.484471 | nearest valid latency median |
| 158 | Jev text | S4-Q04-R1 | B / B | correct | 0.668401 | nearest valid latency median |
| 172 | OpenJev text | S5-Q05-R1 | B / A | incorrect | 1.602863 | nearest valid latency median |
| 198 | OpenJev image | S5-Q07-R1 | C / C | correct | 4.228928 | nearest valid latency median |
In this calculus bank, Jev text scored 127/150 = 84.7%; OpenJev text, image and combined input scored 52.7%, 48.0% and 55.3%. No requests failed. These observed scores describe this fixed question bank and deployment, not a general ranking.
Decision scenarios — five sets
50 base questions × 3 rounds × 4 groups = 600 attempts. Each API’s fixed denominator is 150, including failures.
| API / input | Correct / 150 | Accuracy | Valid accuracy | Failures | HTTP mean s | P50 s | P95 s |
|---|---|---|---|---|---|---|---|
| Jev text | 122/150 | 81.3% | 83.6% (122/146) | 4 | 1.148 | 0.725 | 2.194 |
| OpenJev text | 93/150 | 62.0% | 62.0% (93/150) | 0 | 2.332 | 2.204 | 2.888 |
| OpenJev image | 90/150 | 60.0% | 60.0% (90/150) | 0 | 4.815 | 4.780 | 4.963 |
| OpenJev text + image | 105/150 | 70.0% | 70.0% (105/150) | 0 | 5.875 | 5.735 | 6.422 |
Per-round data
| API / input | Round 1 / 50 | Round 2 / 50 | Round 3 / 50 |
|---|---|---|---|
| Jev text | 39/50 | 43/50 | 40/50 |
| OpenJev text | 32/50 | 30/50 | 31/50 |
| OpenJev image | 30/50 | 32/50 | 28/50 |
| OpenJev text + image | 37/50 | 35/50 | 33/50 |
A question shared by all four APIs
D1-Q08-R1 School poster. One person; hands-on steps cannot overlap or pause. Alone phases can overlap anything. Design: 2m hands. Print: 1m hands, then 5m alone; after Design fully finish. Frame: 4m hands. Mount: 2m hands; after Print + Frame fully finish. Tidy: 2m hands. Follow each listed order, waiting if needed. Which finishes all work earliest? A. Design → Tidy → Print → Frame → Mount B. Design → Print → Frame → Mount → Tidy C. Tidy → Design → Print → Frame → Mount D. Design → Print → Tidy → Frame → Mount Answer key: D
Selected attempt records
| Global # | API / input | Question | Choice / gold | Result | HTTP seconds | Why included |
|---|---|---|---|---|---|---|
| 1 | Jev text | D1-Q08-R1 | B / D | incorrect | 0.937435 | first shared question |
| 2 | OpenJev text | D1-Q08-R1 | B / D | incorrect | 3.942378 | first shared question |
| 3 | OpenJev image | D1-Q08-R1 | B / D | incorrect | 5.378662 | first shared question |
| 4 | OpenJev text + image | D1-Q08-R1 | B / D | incorrect | 7.111950 | first shared question |
| 78 | Jev text | D2-Q03-R1 | — / A | fetch failed [UND_ERR_CONNECT_TIMEOUT] | 10.380249 | all failed attempts |
| 108 | OpenJev text | D2-Q02-R1 | D / D | correct | 2.203423 | nearest valid latency median |
| 150 | OpenJev image | D4-Q10-R1 | C / D | incorrect | 4.780264 | nearest valid latency median |
| 152 | Jev text | D4-Q10-R1 | D / D | correct | 0.716782 | nearest valid latency median |
| 263 | OpenJev text + image | D5-Q09-R2 | B / B | correct | 5.736382 | nearest valid latency median |
| 568 | Jev text | D2-Q06-R3 | — / B | fetch failed [UND_ERR_CONNECT_TIMEOUT] | 10.361398 | all failed attempts |
| 587 | Jev text | D1-Q08-R3 | — / B | fetch failed [UND_ERR_CONNECT_TIMEOUT] | 10.368178 | all failed attempts |
| 600 | Jev text | D3-Q02-R3 | — / D | fetch failed [UND_ERR_CONNECT_TIMEOUT] | 10.385349 | all failed attempts |
Four Jev calls ended in connection timeouts (UND_ERR_CONNECT_TIMEOUT), with no HTTP response. They count as unsuccessful attempts in 122/150 = 81.3%; valid-answer accuracy is 122/146 = 83.6%. These failures do not prove a reasoning error. No retries were substituted. OpenJev combined input scored 105/150 = 70.0%, with a mean HTTP duration of 5.875 seconds in this local deployment.
Interpretation and limits
The two banks are not pooled into one accuracy or timing score. Repeated option orders are not independent questions. There was no external difficulty calibration. The original evaluator clicked the returned option using Playwright; it did not test independent coordinate grounding. Cloud Jev and local OpenJev latency are not a controlled hardware comparison. The original screenshot sizes also differ from the current browser-rendered cards.
Publication and provenance
Only an explicit field allowlist is exported: API group, question IDs, selected/gold answers, correctness, HTTP status/error, UTC timestamps and stage durations. Credentials, Authorization headers, API endpoint URLs, raw request/response bodies, machine paths and internal document links are excluded. The historical measurements are static public examples; visitors’ new results remain in their browser.
The downloadable JSON preserves exact aggregate values and provenance hashes. The CSV preserves original UTC timestamps and all selected timing stages. Full-run source files are not included in this excerpt. Mermaid source for the four charts is in the Markdown download.