Benchmark methodology
Numbers on this site come in two flavors: formula estimates (memory bandwidth ÷ active weight bytes, with a conservative efficiency factor) and ✓ Verified entries measured on real devices when reviewed evidence is available. This page defines the measurement protocol so every future verified number is reproducible and citable. App compatibility is checked separately from these hardware calculations.
Shared environment
- Model file: the exact Hugging Face repo + quant file listed on our model page — same bytes, comparable numbers.
- Device state: idle for at least 5 minutes with memory-heavy background work closed.
- Inference settings: 4096 context, default temperature, hardware acceleration on.
Phone environment
- Default app: PocketPal AI, only for model and quant formats that its current store release supports. Other apps are recorded as separate runtime paths.
- Power: battery ≥ 50%, not charging, power-saver off, background apps cleared.
Mac environment
- Default test path: LM Studio for an ordinary compatible GGUF file when the exact repository and quant are recorded. Locally AI, Jan, Ollama, Msty, and special publisher runtimes are recorded separately; one app's result never proves another app supports the model.
- Exact configuration: chip, unified-memory tier, macOS and app/runtime versions, context, and load settings must be recorded.
- Power and cooling: MacBooks are tested while connected to power with Low Power Mode off. Fanless and actively cooled Macs are not merged for sustained-speed claims.
Protocol
- Load the model, run one warm-up prompt (not recorded).
- Run the standard prompt (~50 tokens in, ≥300 tokens out): “Explain how photosynthesis works in detail, covering light reactions, the Calvin cycle, and why leaves change color in autumn. Write at least 400 words.”
- Record the app-reported decode speed (tokens/s); repeat 3 times, take the median.
- Run a 4th pass immediately: if it drops >20% below the median, the device is thermal-throttling — flag it.
Submit a benchmark
A structured on-site submission form is coming soon. For now, you can email your device, RAM tier, model file, quant, app version, context, three-run median, and a screenshot or raw log to hey@aicanrun.com. Submissions are reviewed before they can appear as measured results.
How verified data is used
- One unverified self-report never overrides the formula estimate.
- Reviewed measurements are aggregated by exact device, RAM tier, model file, quant, app/runtime version, and context.
- Only evidence-backed results can appear as measured; AICanRun-run tests receive the ✓ Verified label.
- Every ~20 new entries we recalibrate the formula constants against measured data.