In the age of artificial intelligence, bigger often seems better. Tech companies frequently tout the massive scale and capability of their latest large language models (LLMs), capable of everything from composing poetry to solving complex scientific problems. However, a new report from the United Nations Educational, Scientific and Cultural Organization (UNESCO) suggests that we might need to rethink that mindset.
According to the 35-page report titled “Smarter, Smaller, Stronger: Resource-Efficient Generative AI & the Future of Digital Transformation,” modest adjustments in how we build and use AI could result in significant energy savings, potentially transforming the sustainability landscape of digital innovation.
Let’s explore the report’s key findings and how they can reshape the future of AI.
1. Why Small Models Matter More Than Ever
One of UNESCO’s central recommendations is clear: use smaller AI models tailored to specific tasks instead of large, general-purpose systems. Not only are these smaller models just as accurate for their intended use cases, but they also consume up to 90% less energy.
Currently, users often rely on massive models trained to do everything — from summarizing articles to generating code. While this all-in-one convenience is appealing, it’s incredibly inefficient. In contrast, using task-specific small models — for example, one trained specifically for translation or customer service queries — matches the model’s strengths with the job at hand, significantly reducing energy consumption and computational demand.
Smaller models are also faster, more cost-effective, and easier to deploy in regions with limited internet connectivity or processing resources. This makes them especially valuable in low-resource settings or organizations with strict cost constraints.
2. Less Talk, Less Energy: The Case for Concise Prompts
Another surprising insight from the report is how much energy can be saved simply by using shorter prompts and generating shorter responses.
LLMs process input text word by word. The more words you use, the more calculations the model has to perform — and each of those calculations requires GPU processing power. Trimming excessive pleasantries or redundant information from your prompts can reduce energy use by over 50% and lower the financial cost of operating these systems.
As Mel Morris, CEO of Corpora.ai, puts it, “The model understands your intent. It doesn’t need pleasantries… but it has to process those extra words anyway, and that costs compute time.”
That said, brevity has its limits. Some tasks require detailed context for accuracy, and too short a prompt may compromise performance. As AI experts point out, the goal is to remove redundancy, not clarity. Smarter prompts, not just shorter ones, are the key.
3. Compress to Conserve: The Power of Model Shrinking
The third strategy recommended by UNESCO involves model compression techniques, such as quantization and pruning. These methods reduce the size and complexity of AI models without significantly affecting their performance. According to the report, effective compression can result in energy savings of up to 44%.
Compressed models take up less space, require less memory, and operate faster. But there’s a trade-off: too much compression can reduce the model’s accuracy or ability to reason through nuanced queries.
As Wyatt Mayham of Northwest AI Consulting cautions, “Overly aggressive pruning or quantization can lead to a drop in accuracy, logical reasoning ability, or nuance, which might make the model unsuitable for its intended purpose.” Implementing these techniques effectively also requires technical expertise and experimentation.
4. Understanding Why Large Models Use So Much Power
At the heart of the energy problem lies model size. Large models contain billions of parameters — the mathematical components that define how a model processes and generates language. Every time an LLM answers a question, it performs massive calculations across all those parameters.
“More parameters mean more calculations,” says Mayham, “which require more GPU power and energy.” It’s similar to how a V8 engine burns more fuel than a four-cylinder one, even when idling.
In contrast, smaller models need fewer parameters and less memory, which means less GPU throughput, faster performance, and lower energy costs. These efficiencies are not only beneficial to the environment but also reduce operational costs for businesses deploying AI tools.
5. The Future of Energy-Efficient AI: A Balanced Approach
While the advantages of smaller models and efficient prompts are clear, experts warn against relying on a single strategy. The most effective path to sustainable AI combines multiple approaches:
-
Use smaller, fine-tuned models for specific tasks
-
Apply compression techniques without compromising performance
-
Write concise, optimized prompts
-
Utilize better hardware and cache common responses
-
Avoid using general-purpose LLMs when a traditional algorithm or smaller model will do
“Even if it is tempting to always throw LLMs at every problem,” said Axel Abulafia of CloudX, “solutions should go from simple to complex. Start with traditional algorithms, move to smaller models, and only use large models when absolutely necessary.”
6. Striking the Balance: Quality, Performance, and Sustainability
The conversation around AI’s environmental impact is still evolving. While large models have captured much of the public’s attention for their impressive capabilities, the future may belong to smaller, more agile, and energy-efficient AI systems.
Achieving this balance between performance and sustainability will require collaboration across the AI ecosystem — from developers and engineers to businesses and policy-makers. As UNESCO’s report makes clear, the question is no longer whether we can make AI greener, but how quickly we can adapt to build smarter and more sustainable digital solutions.
Final Thought: In the world of AI, bigger doesn’t always mean better. Sometimes, being smarter means thinking smaller — and in doing so, we stand to gain a more sustainable and accessible future for artificial intelligence.


