Benchmark methodology

Numbers on this site come in two flavors: formula estimates (memory bandwidth ÷ active weight bytes, with a conservative efficiency factor) and ✓ Verified entries measured on real devices when reviewed evidence is available. This page defines the measurement protocol so every future verified number is reproducible and citable. App compatibility is checked separately from these hardware calculations.

Shared environment

Phone environment

Mac environment

Protocol

  1. Load the model, run one warm-up prompt (not recorded).
  2. Run the standard prompt (~50 tokens in, ≥300 tokens out): “Explain how photosynthesis works in detail, covering light reactions, the Calvin cycle, and why leaves change color in autumn. Write at least 400 words.”
  3. Record the app-reported decode speed (tokens/s); repeat 3 times, take the median.
  4. Run a 4th pass immediately: if it drops >20% below the median, the device is thermal-throttling — flag it.

Submit a benchmark

A structured on-site submission form is coming soon. For now, you can email your device, RAM tier, model file, quant, app version, context, three-run median, and a screenshot or raw log to hey@aicanrun.com. Submissions are reviewed before they can appear as measured results.

How verified data is used