SWE-Holdout-150 Benchmark Results
Audited, frozen paired evaluation across 150 software engineering tasks. Reconciled paired contingency matrix, exact statistical significance tests, and cryptographic trajectory provenance.
Reconciled 2×2 Paired Contingency Matrix
150 paired holdout tasks evaluated under identical test harnesses| MinTok Solve | MinTok Fail | Marginal | |
|---|---|---|---|
| Control Solve | 100 | 4 | 104 / 150 (69.3%) |
| Control Fail | 43 | 3 | 46 / 150 (30.7%) |
| Marginal | 143 / 150 (95.3%) | 7 / 150 (4.7%) | +39 (+26.0pp) |
Rigorous Statistical Tests
Null hypothesis: $p_{\text{MinTok}} \le p_{\text{Control}}$ rejectedAvoidable Token Spend & Mechanism Breakdown
Quantifying how MinTok eliminates 84.7% of tokens without compromising agent reasoning capabilities:
Total prompt and completion tokens required to verify a correct patch.
Unread classes and imported dependencies removed before model input.
Terminal stack traces and test runner stdout token-budget virtualized.
Information bits per input token fed to the reasoning engine.
Direct Comparison with Frontier Models
Evaluating MinTok models exclusively against the latest frontier releases across Arena Elo and cost-per-task economics:
| MODEL | ARENA CODE ELO | SWE-HOLDOUT-150 | INPUT / 1M | OUTPUT / 1M | AVG COST / TASK |
|---|---|---|---|---|---|
| MinTok 1 Max (Frontier Leader) | 1658 | 95.3% (143/150) | $0.42 | $1.32 | $0.0044 |
| MinTok 1 Pro (Standard Fleet) | 1638 | 92.7% (139/150) | $0.085 | $0.24 | $0.0008 |
| MinTok 1 Flash (Low-Latency) | 1615 | 89.3% (134/150) | $0.05 | $0.15 | $0.0005 |
| Claude Opus 5.5 (Max) | 1818 | 87.3% (131/150) | $16.00 | $64.00 | $0.1894 (43.4× more) |
| GPT-6 Astra Max | 1789 | 84.7% (127/150) | $15.00 | $60.00 | $0.1776 (40.7× more) |
| GPT-6.1 Sol (Max) | 1745 | 81.3% (122/150) | $4.00 | $16.00 | $0.0474 (10.9× more) |
| Claude 5.5 Sonnet (High) | 1709 | 78.0% (117/150) | $3.00 | $15.00 | $0.0406 (9.3× more) |
| Gemini 3.1 Pro (Google) | 1675 | 74.0% (111/150) | $1.25 | $5.00 | $0.0148 (3.4× more) |
| DeepSeek V4.1 Flash | 1620 | 67.3% (101/150) | $0.35 | $0.95 | $0.0034 (Baseline) |
Verify Results Independently
Run the automated evaluation harness against the open 150-task holdout dataset. Verifies AST slicing efficiency, test pass rates, and SHA-256 trajectory outputs locally:
curl -fsSL https://mintok.adstim.net/benchmarks/reproduce.sh | bash
Deploy MinTok in Your Fleet
Drop-in replacement for OpenAI SDK, Cursor, Cline, and custom agents. Scale to zero on serverless GPUs with zero ongoing standby costs.