When we talk about generative artificial intelligence, the conversation usually focuses on which model is smarter, which one writes better, or which one programs faster. But there is a factor that rarely enters the equation: the energy cost behind every response. And that cost, dear reader, is what is causing your subscription to rise in price like foam.
ChatGPT and Claude represent two different approaches to the same technological race. OpenAI, with its GPT-4 model, has bet on a generalist and massive approach, with an infrastructure that consumes industrial amounts of electricity to serve hundreds of millions of users. Anthropic, for its part, has followed a different philosophy with Claude, focusing on “constitutional safety” and a more controlled use of inference.
But here is the trap: both models face the same underlying problem. Their data centers need electricity, and that electricity is increasingly expensive. The difference is how they are managing that problem. OpenAI has opted to raise prices more overtly, while Anthropic has applied a more subtle strategy: keeping the nominal price but inflating consumption of tokens through changes in the tokenizer, as we have seen with Opus 4.7.
The data is clear: the cost per million input tokens in Claude is about 5 dollars, and about 25 dollars for output. But as we have documented, the new tokenizer of Opus 4.7 can inflate consumption by up to 35% for exactly the same text. That means that, in practice, the effective cost for the user has risen without the official price having changed by a single cent. ChatGPT follows a similar logic with its own adjustments.
And here is what nobody tells you: the real problem is not which model is better, but that both are building houses of cards on an energy pillar that is cracking. The difference between ChatGPT and Claude is blurring in the face of the common problem: the data centers that feed them are sucking up electricity as if there were no tomorrow, and that tomorrow has already arrived with the bill in hand.
Environmental impact of AI: The elephant in the room that nobody wants to see
Generative artificial intelligence is leaving a carbon footprint that few are willing to acknowledge. When you ask ChatGPT to write you an email, or Claude to review a contract, you are activating physical processes that consume energy and generate heat. And that heat has to be dissipated. And that energy has to be produced. And that production, in most cases, still depends on fossil fuels.
The environmental impact of AI is not an anecdote, it is a systemic problem. The data centers that host these models are responsible for approximately 1-2% of global greenhouse gas emissions, and this figure is expected to grow exponentially. The International Energy Agency estimates that the electricity consumption of data centers will double between 2022 and 2026, going from 460 TWh to more than 1,000 TWh. That is the energy bill we are paying with our tokens.
But the problem goes beyond direct electrical consumption. Data centers need water to cool down, and in many regions of the world, water is a scarce resource. It is estimated that a medium-sized data center can consume up to 600,000 liters of water per day to keep its cooling systems running. That is water that could go to crops, to human consumption, to ecosystems that are already dry.
What really worries me is how this story is being told. Big tech companies have created a narrative of “green AI” that is, at best, a distortion of reality. They talk about efficiency, renewables, carbon offsetting. And yes, it is true that some companies are making significant investments in clean energy. But it is also true that the growth in consumption is outpacing any gain in efficiency. It is the same trap as with cars: we make them more efficient, but since there are more of them, in the end we burn more fuel.
A fact that should make you think: training a model like GPT-3 emitted approximately 552 tons of CO2. That is the equivalent of 123 cars driving for a year. And current models are much, much larger. GPT-4 is estimated to have emitted between 5,000 and 10,000 tons during its training. We are talking about each new model being a small industrial city running at full capacity.
And the worst part is that training is only the tip of the iceberg. Inference —that is, every damn time you use the model— involves an energy consumption that is already surpassing training. In 2023, 60% of the energy consumption of AI data centers went to inference. In 2025, that figure is already approaching 80%. Your daily use of the assistant is leaving a carbon footprint equivalent to that of a small appliance left on non-stop.
Evolution of AI inference costs: From euphoria to reality
Inference is that magical process that occurs when you ask the AI a question and it returns an answer to you. It seems simple, almost instantaneous. But behind that apparent simplicity there is an army of servers burning electricity to process your tokens, interpret your intent, and generate a coherent response.
The cost of inference has been the goose that lays the golden eggs for the tech industry. For years, companies have been burning investors’ money to offer you tokens practically for free. The strategy was clear: hook you, create dependency, and when you can no longer live without it, raise the price. It is the same move Amazon made with free shipping or Spotify with streaming music. But there is a crucial difference: Amazon and Spotify do not have to feed their services with industrial amounts of increasingly expensive electricity.
The data is overwhelming. The inference cost of a model like GPT-4 is estimated at about 4 cents for every 1,000 tokens processed. That does not sound like much, until you do the math. If you have a heavy user who consumes 100,000 tokens a day —which is not at all unreasonable for a developer—, the monthly cost for the company is about 120 dollars. But the user only pays 20. The difference is covered by investors, and that party is coming to an end.
Anthropic, with its Opus 4.7, has been especially transparent —perhaps too much— about this problem. Its gross margin was -94% in 2024. They were losing almost two dollars for every dollar they brought in. In 2025, after a series of adjustments —weekly limits, new tokenizers, blocking third-party tools— they have managed to raise it to 40%. But it is still not enough. They are in a race against the clock to make inference profitable before investors get tired of putting money in.
And here is the trap that makes this evolution so perverse: the cost of inference depends not only on the efficiency of the models, but also on the price of energy. And energy is rising. As we have seen, expensive oil strains the entire system, natural gas skyrockets, and the power companies take the opportunity to raise the price of the megawatt hour. That means that, even if the models are more efficient, if energy is more expensive, the cost per token does not go down.
In 2023, the inference cost of a large language model was about 15 dollars per million tokens. In 2025, thanks to improvements in efficiency, that cost was reduced to about 5-8 dollars. But in 2026, with energy through the roof, companies are estimating that inference costs will exceed 10 dollars per million tokens again. It is the pendulum of efficiency colliding with the reality of a world with finite resources and an electrical grid that cannot keep up.
When we talk about energy efficiency in data centers, we’re not just talking about a server consuming less electricity. We’re talking about a complete ecosystem where every watt of energy has to be managed, cooled, and optimized so the machine doesn’t melt from the heat.
The metric that measures this is PUE (Power Usage Effectiveness). It’s a ratio that compares the total energy consumed by a data center with the energy consumed exclusively by the servers. A PUE of 1.0 would be perfect: all the energy goes to the servers, none to cooling or lighting. The reality is that most data centers have a PUE between 1.2 and 1.6, which means that between 20% and 60% of the energy is spent keeping the machines cool. And with generative AI, that cost skyrockets.
AI data centers are especially voracious because training and inference chips generate far more heat than traditional servers. A GPU like the NVIDIA H100, which is what most AI data centers use, can consume up to 700 watts in operation. That’s the equivalent of having an electric heater running permanently. And in a data center there are thousands of these GPUs running at the same time.
Cooling has become the second-largest operating cost of a data center, only behind the electricity for the servers. And it’s not a problem that can be solved simply by adding more fans. The more heat you generate, the more complex the cooling systems need to be. Some data centers are using direct liquid cooling, which consists of running water through the chips themselves to extract heat more efficiently. But that requires additional infrastructure, pipes, pumps, and energy consumption that add to the final bill.
A fact few take into account: for every watt you consume to run an AI model, you need between 0.5 and 1 additional watt to cool the equipment. That means the real energy cost of each token is not only what the server consumes, but also what the cooling systems consume. We’re paying double for the same energy.
And the situation gets worse because AI data centers need to be close to energy sources and high-speed telecommunications networks, which limits them to specific areas. In Virginia, which is the largest data center hub in the world, the electricity demand for data centers already exceeds 25% of the region’s total consumption. The local power grid is on the verge of collapse and the waiting lines for new interconnections exceed four years. It’s not that tech companies don’t want to expand, it’s that they physically can’t.
Generative AI and electricity consumption: The bill that didn’t show up in investors’ slides
Generative AI has been sold as a technological revolution comparable to the invention of the internet. And in a way, it is. But what nobody tells you in investors’ presentations is that this revolution needs an energy infrastructure that doesn’t exist and that is straining the electrical systems of half the world.
When we talk about generative AI and electricity consumption, the data is overwhelming. The most conservative estimates place the electricity consumption of generative AI at the equivalent of a country like Spain. And the most aggressive estimates, the ones that take exponential growth into account, say that by 2030 AI could consume as much as the entire United Kingdom. This is not an exaggeration. It’s the projection that energy analysts are making when they take the pace of data center construction into account.
The explosion of generative AI has caught the energy industry with its pants down. In 2023, the electricity consumption of AI data centers was around 25 TWh. In 2026, that figure has already surpassed 100 TWh. That’s 300% growth in three years. No other sector has grown so fast in recent history, not even electric vehicles.
And what’s worse: generative AI is not flexible consumption. You can’t tell an AI model “don’t think so much right now because it’s peak time on the power grid.” When millions of users use ChatGPT at the same time, the demand is immediate and inelastic. Power grids can’t adapt to this type of peak consumption without having backup capacity, and that backup capacity is provided by, of course, gas and coal plants and, in some cases, nuclear plants.
A concrete example: the integration of AI into Google searches, which they announced with great fanfare, multiplies electricity consumption by 10 compared to a traditional search. That means that if all Google searches had an AI component, the company’s electricity consumption would multiply by 5. And Google isn’t the only one. Microsoft, Amazon, Meta, they’re all integrating AI into their core products.
The result is a perfect storm of electricity demand that is colliding with an energy supply that can’t grow at the same pace. Renewables are a solution, but they can’t provide the stability that data centers need 24/7. Nuclear energy is stable, but building a plant takes 10 years. Gas is flexible, but it’s expensive and emits carbon. Coal is cheap, but it’s an environmental disaster. And while we decide, the price of electricity keeps going up and up.
Comparing cloud service costs: The art of charging you more for doing less
Comparing the costs of cloud services that run AI is like comparing the price of electric cars: they all tell you they’re cheaper to maintain, but then you see the battery bill and your jaw drops. AWS, Google Cloud, and Azure clouds have turned the cost of inference into a sleight-of-hand game where the numbers dance however they please.
Let’s start with the basics: the hourly cost of a GPU in the cloud. AWS charges you about $8 an hour for an NVIDIA H100. Azure about $9.50. Google Cloud about $8.20. At first glance it seems like they’re all competing. But the fine print is where the dance begins: data egress costs, storage, commitment credits, volume discounts…
A developer running an AI model in the cloud doesn’t just pay for the GPU. They pay for the data that comes in, the data that goes out, storage, API calls, inter-region transfers. The real cost of inference in the cloud can be between 2 and 5 times higher than the nominal cost of the GPU. It’s like when you go to the dealership and the car costs 20,000 but by the time you leave with insurance, taxes, maintenance, and fuel, the figure has gone up to 30,000.
GitHub Copilot is the perfect example of how to inflate costs without it looking like a price hike. Before, a user could access Opus 4.6 for 3 premium credits. Now, Opus 4.7 costs 7.5 credits. That’s a 150% increase in the effective cost for the user. The official excuse is that the model “thinks more” and that the tokenizer is less efficient. But as we’ve seen, the tokenizer inflates consumption by between 0% and 35%. Even assuming the worst-case scenario, 3 credits for 1.35 is 4 credits. Where do the other 3.5 credits come from? That’s pure and simple commission, an overcharge with no technical justification.
And it’s not just GitHub. Anthropic’s subscriptions for intensive API use have already started including token limits that didn’t exist before. They sell you a $200-a-month subscription with a weekly limit, and if you go over, you enter a pay-per-use system that can hit you with $2,000 in a day without warning. The real cost of inference in the cloud has become an unpredictable money pit that empties developers’ wallets.
The overall trend is clear: cloud services are shifting the AI bill from investors to end users. For years, big tech companies have been burning money to gain market share. Now that they’ve got millions of users hooked and energy prices are through the roof, they’re applying the same model as utility companies: they charge you for actual consumption, with prices that rise when demand tightens.
Cooling systems in data centers: The war against heat that nobody sees
Cooling is the Achilles’ heel of artificial intelligence. Every chip running an AI model generates as much heat as a car radiator. And when you have thousands of chips in the same room, the heat they generate is enough to change the microclimate of the area where the data center is located.
Cooling systems have had to evolve at the speed of light to keep up with AI demand. Traditional air conditioning systems are no longer enough. We’re seeing solutions ranging from direct liquid cooling to total immersion of servers in dielectric fluids, where the chips are literally submerged in a bath of special oil that absorbs the heat and is then cooled in a heat exchanger.
The cost of these cooling systems is astronomical. A next-generation data center can spend between 30% and 50% of its energy budget on cooling alone. And that’s without counting the initial investment in pipes, pumps, heat exchangers, and control systems. The most advanced data centers are using seawater or river water cooling in places where water is naturally available. But that severely limits where you can build, because not every site has a river nearby.
The data is shocking: a medium-sized data center consumes 600,000 liters of water per day for cooling. That’s the equivalent of the water consumption of 10,000 people. And AI data centers are anything but small. Microsoft’s data center in Iowa, for example, consumes 1.5 million liters of water per day. That’s the equivalent of 1,000 Olympic swimming pools per year. We are diverting water from ecosystems that are already dry to cool machines that write emails for us.
The problem is compounded because the heat generated by AI chips is so concentrated that traditional cooling systems cannot dissipate it efficiently. An H100 GPU generates between 500 and 700 watts of heat in a space the size of the palm of a hand. That’s a thermal density that no air conditioning system can handle without consuming industrial amounts of energy. That’s why data centers are moving to liquid cooling, which can dissipate heat much more efficiently.
But liquid cooling has its own problems: it requires special infrastructure, energy-consuming pumping systems, pipes that can leak, and much more complex maintenance. And despite all this, the energy efficiency of data centers is getting worse because chips are increasingly powerful and generate more heat per square centimeter. We are winning the performance battle but losing the efficiency war.
The Future of AI Infrastructure: A Horizon of Tension and Contradiction
The future of AI infrastructure is full of contradictions that would make any economist pale. On one hand, everyone wants more AI. On the other hand, nobody wants to pay the energy cost that entails.
The growth projections are simply unsustainable. Demand for AI computing capacity is expected to grow 30% annually over the next decade. By 2035, AI data centers could be consuming 4.4% of all global electricity. That’s more than the combined consumption of countries like Germany, France, and the United Kingdom. It’s not that AI is going to consume a lot of electricity, it’s that it’s going to consume all the electricity we can produce and a little more that we don’t have.
And the problem is not just one of quantity, but also of geographic and temporal distribution. AI data centers need to be near stable energy sources and high-speed communication networks. That means they concentrate in a few regions of the world: Virginia, California, Texas, Ireland, Singapore. These regions already have strained power grids, and the arrival of these data centers is causing delays in connecting other projects, including hospitals and residential neighborhoods.
The bottlenecks are not just energy-related. The chip supply chain, the availability of transformers, the materials for cooling systems—everything is at its limit. We are in a situation where we cannot build new data centers because there isn’t enough transformer manufacturing capacity, nor enough high-end chips, nor enough energy transmission capacity to power all of this.
Big tech companies are making decisions that five years ago would have seemed like science fiction. Microsoft has signed agreements to restart nuclear power plants. Google is investing in small modular reactors. Amazon is buying large-scale solar energy but has also increased its dependence on natural gas. The contradiction is total: the companies that most proclaim their commitment to the environment are the ones keeping coal and gas alive.
AI infrastructure is becoming a battleground between the need for cheap energy and the obligation to reduce emissions. And for now, cheap energy is winning. We are seeing how coal is once again becoming a viable option in the United States, how natural gas is establishing itself as the transition fuel, how nuclear plants are being reconsidered even in countries that had ruled them out. AI is killing climate change in real time, and it’s doing so in the name of progress.
The future, I fear, is a two-speed world. On one hand, those who can afford the energy will have access to the best AI, the fastest models, cutting-edge research. On the other, those who can’t afford it will be left with crippled models, token limits, and bills that will make them cry. AI is not going to democratize knowledge; it’s going to perpetuate the planet’s energy inequalities. And that, dear reader, is the story nobody is telling on the evening news.
And so, between expensive tokens, models that barely chug along, data centers that suck electricity like an octopus, and energy that won’t stop rising, we find ourselves in 2026 with a reality that nobody wants to see. The dream of cheap and abundant AI is fading amid columns of smoke from coal plants and the incessant hum of cooling systems struggling to keep cool a monster that won’t stop growing. The question is not whether AI is going to change the world, but who is going to pay the bill for that change. Spoiler: it won’t be those who are getting rich off it.