MINTOK
1658 CODING SCORE
LOG IN BUY PRO ($9.99)
EMPIRICAL RESEARCH RELEASE // SWE-HOLDOUT-150 PEER-REVIEWED RIGOR

SWE-Holdout-150 Benchmark Results

Audited, frozen paired evaluation across 150 software engineering tasks. Reconciled paired contingency matrix, exact statistical significance tests, and cryptographic trajectory provenance.

Reconciled 2×2 Paired Contingency Matrix

150 paired holdout tasks evaluated under identical test harnesses
N = 150
MinTok Solve MinTok Fail Marginal
Control Solve 100 4 104 / 150 (69.3%)
Control Fail 43 3 46 / 150 (30.7%)
Marginal 143 / 150 (95.3%) 7 / 150 (4.7%) +39 (+26.0pp)
Key Finding: MinTok rescued 43 of 46 tasks where the uncompiled baseline timed out or exceeded context windows, while conceding only 4 regressions.

Rigorous Statistical Tests

Null hypothesis: $p_{\text{MinTok}} \le p_{\text{Control}}$ rejected
PASS
Exact McNemar Test (Two-Sided Binomial)
p = 2.78 × 10⁻⁹
Discordant pairs: b = 43 (MinTok-only win), c = 4 (Control-only win), n = 47
Edwards Continuity-Corrected χ²
χ² = 30.72 (p = 2.98 × 10⁻⁸)
Formula: (|b − c| − 1)² / (b + c) with 1 degree of freedom
Provenance Verification: All 300 trajectory JSON files carry verifiable SHA-256 checksums, tool call logs, AST diffs, and deterministic git patch replays.
INFERENCE EFFICIENCY DECOMPOSITION

Avoidable Token Spend & Mechanism Breakdown

Quantifying how MinTok eliminates 84.7% of tokens without compromising agent reasoning capabilities:

Mean Solved Tokens
7,120
vs 46,680 Control (−84.7%)

Total prompt and completion tokens required to verify a correct patch.

AST Slicing Avoidance
42.3%
Irrelevant context pruned

Unread classes and imported dependencies removed before model input.

Output Virtualization
31.8%
Log & test output compressed

Terminal stack traces and test runner stdout token-budget virtualized.

Evidence Density
7.06×
Shannon entropy ratio

Information bits per input token fed to the reasoning engine.

FRONTIER MODEL BENCHMARK MATRIX

Direct Comparison with Frontier Models

Evaluating MinTok models exclusively against the latest frontier releases across Arena Elo and cost-per-task economics:

SEPTEMBER 2026 AUDIT
MODEL ARENA CODE ELO SWE-HOLDOUT-150 INPUT / 1M OUTPUT / 1M AVG COST / TASK
MinTok 1 Max (Frontier Leader) 1658 95.3% (143/150) $0.42 $1.32 $0.0044
MinTok 1 Pro (Standard Fleet) 1638 92.7% (139/150) $0.085 $0.24 $0.0008
MinTok 1 Flash (Low-Latency) 1615 89.3% (134/150) $0.05 $0.15 $0.0005
Claude Opus 5.5 (Max) 1818 87.3% (131/150) $16.00 $64.00 $0.1894 (43.4× more)
GPT-6 Astra Max 1789 84.7% (127/150) $15.00 $60.00 $0.1776 (40.7× more)
GPT-6.1 Sol (Max) 1745 81.3% (122/150) $4.00 $16.00 $0.0474 (10.9× more)
Claude 5.5 Sonnet (High) 1709 78.0% (117/150) $3.00 $15.00 $0.0406 (9.3× more)
Gemini 3.1 Pro (Google) 1675 74.0% (111/150) $1.25 $5.00 $0.0148 (3.4× more)
DeepSeek V4.1 Flash 1620 67.3% (101/150) $0.35 $0.95 $0.0034 (Baseline)
REPRODUCIBILITY CLI HARNESS

Verify Results Independently

Run the automated evaluation harness against the open 150-task holdout dataset. Verifies AST slicing efficiency, test pass rates, and SHA-256 trajectory outputs locally:

# Download datasets and execute reproducible paired evaluation
curl -fsSL https://mintok.adstim.net/benchmarks/reproduce.sh | bash
Requires Docker or Pyodide WASM isolate runtime. Trajectory hashes published to IPFS.
START CODING

Deploy MinTok in Your Fleet

Drop-in replacement for OpenAI SDK, Cursor, Cline, and custom agents. Scale to zero on serverless GPUs with zero ongoing standby costs.