
Hardware-Software Interplay: Measuring the Energy Cost of Local LLM Inference Across Integrated Testbed Profiles on Consumer Hardware
Aaryan Karlapalem
21/07/2026
As large language models (LLMs) evolve from data centers to local edge devices, understanding the hardware-software interactions of these systems has become essential for creating sustainable forms of AI. This study investigates the influence of host operating systems and machine specifications on the overall energy efficiency of three common on-device LLMs: Gemma 2:2B, Llama 3.2:3B, and Mistral 7B, running locally via Ollama. By measuring both the average power draw and execution time across the 3 most common consumer operating systems of macOS 26, Windows 11, and Debian Linux 13, the 'energy-per-token' cost of local inference can be calculated. The experimental results indicated that Mistral 7B consistently requires the highest energy input, while Gemma 2:2B is far more efficient for shorter tasks. Statistical analysis via a Two-Way ANOVA (p<0.05) confirms that the integrated testbed profiles significantly influence energy outcomes. These findings provide a concrete framework for optimizing local AI deployments to balance performance with environmental sustainability.