LFM2.5-VL 3B

At the default Q4_K_M, 138 of 139 phones have enough estimated usable memory; 138 also reach 3+ tokens/s. On Mac, 54 of 54 exact configurations pass using that quant or a smaller tracked fallback.

HOW TO RUN · OFFICIAL GGUF

Use Liquid AI's official GGUF. Vision also needs the separate Q8 projector (about 0.58GB), so the main model download alone understates working memory. Q4_K_M plus the projector is about 2.26GB before KV cache and runtime overhead.

llama-mtmd-cli -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M --image ./image.jpg
Official model card and current llama.cpp instructions ↗
Params
3.1B
Family
lfm
Released
2026-08
Tags
chat · vision

Downloads & memory by quant

The download is only the weights — add KV cache and runtime to get what your phone actually needs. “Memory-fit” only means it can load; “usable” also requires an estimated 3+ tokens/s. Each row uses that row's quant. ★ = recommended.

QuantDownloadMemory neededPhone reality
Q4_01.6 GB~2.5 GB138 memory-fit · 138 usable
Q4_K_M1.7 GB~2.6 GB138 memory-fit · 138 usable
Q5_K_M1.9 GB~2.8 GB138 memory-fit · 138 usable
Q6_K2.2 GB~3.1 GB138 memory-fit · 138 usable
Q8_02.9 GB~3.9 GB138 memory-fit · 127 usable

LFM2.5-VL 3B at Q4_K_M on 139 phones

This table is fixed to the recommended Q4_K_M, on each phone's largest RAM option. Counts for smaller quants above will differ. All phones are shown, sorted by fit and estimated experience; tap one for the full report.

PhoneChip · RAMNeedsSpeedVerdict

~ = bandwidth-based estimate · ✓ = measured on real hardware

LFM2.5-VL 3B on 54 Mac configurations

Each Mac uses the recommended Q4_K_M at 4K context when it fits, or the largest smaller tracked quant that fits. Each row links to the full report with a macOS-specific setup guide.

MacUnified memoryQuantWorking budgetNeedsEstimated speedVerdict

FAQ

How much RAM do I need to run LFM2.5-VL 3B on a phone?

At Q4_K_M, LFM2.5-VL 3B needs ~2.6GB of usable memory (weights + KV cache + runtime). In practice that means a 6GB+ Android phone or a 6GB+ iPhone — iOS lets a single app use less of its RAM than Android does.

How fast does LFM2.5-VL 3B run on a flagship phone?

On the Galaxy S25 Ultra (Snapdragon 8 Elite) it runs at ~20.3 tokens/s at Q4_K_M, estimated from memory bandwidth. Anything above ~8 tokens/s feels smooth for chat.

Can an iPhone run LFM2.5-VL 3B?

Yes — the iPhone 16 Pro Max runs it at ~15.9 tokens/s at Q4_K_M, using 2.6GB of its ~5.2GB usable memory.

What is the best quantization of LFM2.5-VL 3B for mobile?

Q4_K_M (1.7GB download) is the size/quality sweet spot of the 5 quants available. Total memory needed is ~2.6GB once the KV cache and runtime are counted — the download size alone understates it.

How many phones can run LFM2.5-VL 3B?

138 of the 139 phones we track have enough estimated usable memory at Q4_K_M; 138 also reach our usable-speed threshold of 3 tokens/s, and 106 run smoothly. Memory-fit alone does not mean a good experience.

Which Macs can run LFM2.5-VL 3B?

54 of the 54 exact Mac configurations we track have enough estimated working memory. Each Mac uses the recommended Q4_K_M when it fits, or the largest smaller tracked quant that fits, and separates installed unified memory, the macOS reserve, and estimated speed.

What is LFM2.5-VL 3B good for on a phone?

It's tagged for chat, vision. At 3.1B parameters, it's a fast, lightweight pick for quick tasks.

Similar models

LFM2.5 2.6BLFM2.5 8B-A1BTini Cybersec 8B-A1BSmolLM3 3BLlama 3.2 3B