Local AI can now handle 88.7% of real-world queries, study finds

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

Researchers evaluated 20+ local language models on 1M real-world queries, measuring accuracy, energy, latency, and power. They introduce intelligence per watt (IPW) as a unified metric. Local models answered 88.7% of queries, and IPW improved 5.3x from 2023 to 2025. Local accelerators achieved at least 1.4x better IPW than cloud accelerators on identical models, suggesting local inference can meaningfully redistribute demand from centralized infrastructure.

Local accelerators achieve at least 1.4x lower IPW than cloud accelerators running identical models, revealing significant headroom for local accelerator optimization.

More from this day

2026-09-16