# Timlig — full content
> Field notes from building high-performance systems and teams. Long-form engineering writing on AI infrastructure, sovereign AI in India, GPUs, semiconductors, data center supply chains, and how high-performance teams scale. Written by Anuj Sharma.
Site: https://www.timlig.com
RSS: https://www.timlig.com/rss.xml
Below are the full, unmodified contents of every published post on Timlig, latest first.
---
---
# RSS is live
Source: https://www.timlig.com/posts/rss-is-live/
Published: 2026-05-08
Tags: meta, rss
Timlig now has an RSS feed at [/rss.xml](/rss.xml). Add it to your reader of choice and every new post lands there automatically — no email, no signup, no algorithm.
---
# The Two-Rupee Voice
Source: https://www.timlig.com/posts/two-rupee-voice/
Published: 2026-05-07
Tags: AI, India, BPL, welfare, Sarvam, IndiaAI Mission, ASR, TTS, language models, policy
# The Two-Rupee Voice
Here is a fact about Indian AI in 2026 -
A full voice conversation in Hindi — speech-to-text, a language model that thinks for a second, text-to-speech that talks back, and a translation step in the middle — now costs roughly **₹1.85 to ₹2.05** end-to-end. Call it two rupees. That is about **30 to 60 seconds of MGNREGA wage** under the FY 2025–26 wage notification. It is also one-fifteenth of what the same conversation would have cost on a frontier-vendor API in mid-2024, when GPT-4-class tokenizers were charging Hindi-speakers four to eight tokens per word — versus 1.4 for English — for the exact same meaning. The “token tax.”
The token tax is dead. An Indian-built 2-billion-parameter model called **Sarvam-1**, trained from scratch on 4,096 H100s on Yotta’s Shakti cloud, dropped Hindi token fertility to **1.4–2.1**. **Saaras V3**, an Indian-built ASR system, reaches **19.31% word error rate** on the IndicVoices 10-language test set at **₹30 per hour of audio**. **Bulbul**, an Indian-built TTS system, costs **₹15 per 10,000 characters**. The compute underneath all this is, courtesy of the IndiaAI Mission, **₹65–92 per H100-hour** with a 40% subsidy that drops the effective rate to about **₹40/hour** for empanelled users. A leading commercial Indian neocloud charges around ₹249/hour for the same chip. AWS Mumbai charges roughly ₹330. The Indian government, when it wants a Hindi sentence to come out of a speaker, is paying *somewhere between zero and twelve percent of what AWS would charge*.
This is, in industrial-policy terms, a triumph. It is also — and this is the part I want to spend the next several thousand words on — *the moment the economics of giving every poor Indian a personal AI advisor structurally crossed below the cost of any single welfare program in the country*. And almost nobody is currently buying.
Let me explain.
-----
## 1. The basic problem
The basic problem is that there are, depending on which number you trust, somewhere between **129 million and 234 million** Indians living in poverty, and they have always had an information problem.
NITI Aayog’s January 2024 Discussion Paper on Multidimensional Poverty puts the headcount ratio at **11.28% in 2022–23**, down from 24.85% in 2015–16, with **24.82 crore people exiting** in eight years. UP alone moved 5.94 crore out of MPI poverty; Bihar 3.77 crore; MP 2.30 crore. The UNDP’s Global MPI 2024, which uses the older NFHS-5 round, still classifies **234 million** Indians as multidimensionally poor — the largest national cohort in the world. The Household Consumption Expenditure Survey 2022–23, the first such round since 2011–12, implies a Tendulkar-equivalent monetary poverty of 5–10%. The Government of India, asked under RTI in December 2024 by *Down To Earth* for its current official poverty count, replied — and I am compressing this — *we don’t actually maintain one*.
So: a number, somewhere between 130 and 240 million, mostly concentrated in Bihar, UP, Jharkhand, MP, Odisha, Rajasthan, Assam, Chhattisgarh; speaking, mostly, *not* the languages OpenAI’s tokenizer was optimized for; relying on a frontline-worker network — about a million ASHAs, 1.3 million Anganwadi workers, 50,000 agricultural extension officers — that is structurally undersupplied; and, since the JAM trinity finished its job, fully addressable via Aadhaar and UPI and increasingly via WhatsApp.
What this person needs from an AI is not particularly mysterious. *Is my crop disease bacterial or fungal? When does my PM-KISAN installment land? Is my pregnancy at risk? What does my MGNREGA wage statement actually say? My Aadhaar got delinked from my ration card — what do I do?* The answers to these questions are, mostly, sitting in some government database or Krishi Vigyan Kendra advisory that the user cannot read in any language they speak. The agent-to-farmer ratio in the Indian extension system is **1:650**. The ASHA-to-population ratio is roughly **1:1,000**. The information layer is broken not because the information doesn’t exist but because the last-mile translation of that information into the user’s *spoken Bhojpuri or Maithili or Santhali* is, structurally, missing.
The thing AI is genuinely good at is the spoken-Bhojpuri-to-government-database translation. The thing AI was, until eighteen months ago, far too expensive to do at population scale. *That second thing has changed. The first thing has not.*
-----
## 2. The token tax, and what its death means
I want to dwell on the token tax for a moment, because it is the cleanest example of how the global frontier was, for several years, charging a structural surcharge to people who could not afford it.
A tokenizer is the thing that breaks an input string into the units a language model actually thinks in. Roughly, OpenAI’s `cl100k_base` tokenizer — the GPT-4 / GPT-3.5 era — assigns about **1.4 tokens per word** for English. For Hindi, the same tokenizer assigns roughly **4 to 6 tokens per word**. For Tamil, 7 to 8. For Malayalam, in some samples, **double-digit tokens per word**. The reason is mundane: the tokenizer was trained on a corpus that was overwhelmingly English, so English words got single-token “shortcuts” while Devanagari/Tamil/Malayalam scripts had to be decomposed into characters or sub-character UTF-8 fragments.
What this looks like in practice: if you and I are asking the same question, and you ask in English and I ask in Tamil, *I am paying five times your bill*. For the same meaning. On the same model. At the same per-token price. When OpenAI released GPT-4o’s `o200k_base` in 2024, Microsoft’s announcement noted that the new tokenizer cut Tamil tokens by about 74% and Malayalam by roughly 4×. A tacit admission, if you read it the right way, that the previous billing was structurally weird.
The Indian fix arrived in October 2024. Sarvam-1, a 2-billion-parameter dense decoder trained from scratch by Sarvam AI on 2 trillion tokens via NVIDIA NeMo on 4,096 H100 SXMs in Yotta’s Shakti cloud, dropped Hindi token fertility to **1.4–2.1**. Telugu, Kannada, Tamil, Bengali, Marathi, Gujarati, Punjabi, Malayalam, Odia, English: ten Indic-plus-English languages, each tokenized at roughly the same fertility as English. The model is openly available; the API is rupee-priced; the inference is reportedly four to six times faster than Gemma-2-9B and Llama-3.1-8B on Indic tasks.
The practical consequence is that the *price of one unit of meaning* in Hindi is now structurally similar to the price of one unit of meaning in English. The poor Indian’s linguistic surcharge — let’s call it what it was — has been refunded.
The price-per-meaning collapse compounds with two other things:
The first is **ASR cost**. Sarvam’s published Saaras V3 rate is ₹30 per hour of audio (about $0.36), at 19.31% WER on the 10-language IndicVoices subset. Whisper-large-v3, OpenAI’s flagship Indic-enabled ASR, would cost five to ten times more per hour at retail and would, on rural Bhojpuri women’s voices specifically, be measurably worse — there is an entire AI4Bharat dataset called “Bhojpuri and Hindi Rural Women ASR” because this gap had to be addressed by hand.
The second is **compute cost**. The IndiaAI Mission has onboarded 34,000+ GPUs across 14 empanelled providers — Yotta, E2E Networks, Tata Communications, Jio Platforms, CtrlS, AWS MSPs and others — at a tendered floor of ₹65/GPU-hour, with H100s at ₹92/hour and a 40% subsidy for approved researchers, MSMEs, startups, and government users. Sarvam alone received 4,096 H100s and a ₹98.68 crore subsidy against a ₹246.71 crore project award. The mission’s total approved sanction across 12 foundation-model awardees is over ₹2,000 crore.
Stack the three: tokenizer fertility down ~3–4×; ASR cost down ~5–10×; compute cost down ~5–9× against AWS retail. The unit cost of one Indian-language voice round-trip has fallen by something like **fifteen to forty times** in eighteen months, depending on the language and the workload.
This is what people mean when they say “the price has crossed the threshold.” It is a real number; it is published; it is invoiced.
-----
## 3. The two-rupee voice
Here is the actual unit-economics arithmetic, which I want to walk through carefully because it is the central fact of this essay.
A full voice round-trip — a poor Hindi-speaking farmer asks his phone whether his cotton crop has bollworm, the system transcribes, retrieves, generates, translates, speaks back — has roughly four cost layers:
|Layer |Provider/system |Cost per query|
|----------------------------------------|-----------------------------------------------------|--------------|
|ASR (≤60 sec audio) |Sarvam Saaras V3 at ₹30/hr |₹0.50 |
|LLM inference (~500 in / 200 out tokens)|Sarvam-1 / Llama 3.2 3B on subsidized IndiaAI compute|₹0.05–0.20 |
|TTS (~200 chars output) |Bulbul v2 at ₹15/10K chars |₹0.30 |
|Translation (~500 chars) |IndicTrans2 / Bhashini at ₹20/10K chars |₹1.00 |
|**Total per voice query** | |**₹1.85–2.00**|
Round it: **₹2 per full voice conversation**. Two rupees. About 2.4 US cents.
Compare to:
- **MGNREGA wage**, FY 2025–26, national average roughly **₹370/day** for an 8-hour workday. Per-minute wage: ~₹0.77. Per second: about 1.3 paise. *One AI voice query equals about two and a half minutes of MGNREGA wage.*
- **Mobile data** at ₹8–10/GB retail; a 60-second voice round-trip pushes maybe 100 KB up and 50 KB of compressed audio back. Data cost per query: <₹0.001. Negligible. Already paid for.
- A **Jio Bharat ₹125/month** plan is unlimited voice and minimal data; the connectivity cost per AI query, amortized, is a fraction of a paisa.
- **A two-minute call to a doctor on a paid telemedicine app**: typically ₹50–200. The AI alternative is *25–100× cheaper per interaction*. (It is also, of course, not a doctor. We will get to this.)
Now run it at population scale.
A daily AI advisory — say, **5 voice queries per day** per BPL adult — for India’s 234 million MPI-poor population, at retail ₹2/query, costs:
> 234,000,000 × 5 × ₹2 × 365 = **₹85,410 crore per year**.
That is, by retail accounting, roughly equivalent to one entire MGNREGA budget (~₹86,000 crore for FY 2025–26). It is the *upper bound* — the bill if every Indian below the poverty line consumed a doctor-grade conversational AI five times a day, at full sticker. It is, in welfare-state terms, *not actually that big*.
Now run it at subsidized IndiaAI prices. Compute is the dominant variable cost; everything else (ASR, TTS) scales with the same subsidized GPU pool. Apply a conservative 60% all-in cost reduction on the LLM layer and 40% on the speech layers — a fair mid-range estimate of what empanelled IndiaAI access actually unlocks:
> Effective per-query cost ≈ ₹0.80–1.20 → **₹1/query** as a working number.
>
> 234M × 5 × ₹1 × 365 = **₹42,705 crore/year**.
Now narrow the scope. Say the goal is not five queries per BPL Indian per day but **one structured voice query per BPL household per day**, which is a more realistic adoption curve. India has roughly 50 million MPI-poor households. One query per household per day, at ₹1 effective cost:
> 50M × 1 × ₹1 × 365 = **₹1,825 crore/year**.
That is **less than 3% of the PM-KISAN budget** (₹60,000 crore/year). It is about 2% of MGNREGA. It is one-tenth of what the Ministry of Rural Development spends *only* on the Pradhan Mantri Awaas Yojana – Gramin in a busy year. If you wanted to give every poor household in India a daily, voice-mediated conversation with a personalized scheme/agriculture/health advisor — at scale, in their own language — the upper bound on the bill is *smaller than the line items on the existing welfare receipts that the same household already receives*.
The economics, *for the first time in the history of post-Independence India*, are not the binding constraint.
-----
## 4. The phone is the binding constraint, but only just
It would be very Timlig of me to stop here and say “the economics work, end of story.” The economics work. The phone is the next problem.
Here is what the BPL household actually owns, roughly:
The **Jio Bharat 4G feature phone**, ₹799–1,199, 512 MB RAM, 4 GB storage. It cannot run an SLM on-device. It can consume cloud-mediated voice over a UPI 123Pay-style flow — IVR plus cloud LLM plus TTS streamed back. Tens of millions of these.
The **Redmi A series, Realme C series, Itel/Infinix budget Android**, ₹6,000–9,000, 4 GB RAM, MediaTek Helio G35–G88 or Snapdragon 4 Gen 2. A 1B-parameter SLM at 4-bit quantization, ~600 MB on disk, ~1 GB RAM at runtime, *runs* — at maybe 3–8 tokens per second for short Hindi voice replies. Borderline usable. Not pleasant.
The **mid-tier Xiaomi/Realme/Samsung A-series**, ₹12,000–18,000, Snapdragon 6 Gen 1 / Dimensity 7000-class, 6–8 GB RAM. A 3B model at 4-bit (~2 GB on disk) is plausible. This is the male earner’s phone in many BPL households, often the *household’s only smartphone*.
ASER 2024 finds that **90% of 14–16-year-olds in rural households have a smartphone at home**, but only **31% own one personally**, and the share that uses it for educational purposes is **57%**. NFHS-5 / Data For India: phone *usage* among the poorest women rose from 39% to 67% in three years; among the poorest men, from 74% to 84%. A persistent 10+ point gender gap. The household has access to a smartphone; the woman often does not control it. The model that fits the poorest household’s phone is too small for nuanced dialect understanding, and the model that handles the dialect doesn’t fit the phone.
Which means, today, the credible deployment topology is **not on-device autonomous agent for the poor** — that is a 2027–2028 story, when the median ₹8,000 budget Android crosses 6 GB RAM and Snapdragon 6 Gen 1 — but **cloud-mediated voice**, accessed via:
- **WhatsApp** (78% of Indian smartphones, including BPL), the way Jugalbandi, Farmer.CHAT, Kisan e-Mitra, and ASHABot already deliver. The phone records audio and plays audio. The intelligence is in the cloud, on subsidized H100s.
- **IVR over feature phones**, the way Kilkari (3M+ active maternal mHealth subscribers) already works — upgraded from one-way pre-recorded audio to two-way voice agents.
- **Frontline workers’ phones**, the way ASHABot routes through 869 ASHAs in Udaipur. The worker has a Snapdragon 6, the citizen has a feature phone, the worker mediates.
- **Common Service Centres, kirana stores, Bank Mitras**, the way the JAM trinity already mediates. About 400,000 CSCs, 1.2 million kirana stores, and PM-WANI’s 4,09,111 Wi-Fi hotspots. The infra is there.
The cloud-mediated voice topology, on subsidized IndiaAI compute, *already passes the price test* for any meaningful BPL deployment. The on-device topology will pass it in 2027–28. Either way, the phone stops being the binding constraint by 2028.
-----
## 5. The boring-but-essential second-order thing: the missing payer
This is the part of the essay that earns its keep. It is also the part that makes everything above slightly depressing.
The unit economics work. The model is built. The compute is subsidized. The data exists — IndicVoices is **7,348 hours across 22 languages, 16,237 speakers, 145 districts**, funded for ~₹30 crore by Bhashini, EkStep, and Nilekani Philanthropies. The DPI rails — Aadhaar, UPI, DigiLocker, Account Aggregator, ONDC, Bhashini, BharatNet’s **2.15 lakh gram panchayats** with 4G and 5G coverage in 99.9% of districts — are unusually good. The frontline worker network — a million ASHAs, 1.3 million Anganwadis, 50,000 extension officers — is in place, structurally undersupplied, and ready to be augmented.
What is missing is *the buyer of inference at the back end*.
In every functioning AI economy, somebody pays per query. ChatGPT users pay $20/month. Enterprise customers pay per-token through API contracts. Advertisers pay through Google. The economic loop closes because the consumer of the inference is also the funder of it, or because an advertiser triangulates between them.
For BPL Indians, neither holds. The user cannot pay $20/month for ChatGPT — that is roughly two days of MGNREGA wage. The user is not an advertising target attractive enough to sustain a venture-funded direct-to-consumer model — Karya is a real exception precisely because it pays *into* BPL households rather than monetizing them, but Karya is a labor platform, not an inference platform. The user is, in the language of welfare economics, somebody whose AI consumption produces *positive externalities* (better health, higher agricultural yield, fewer wrongful scheme exclusions) that are not captured by the user’s willingness-to-pay.
In every other welfare program, India has solved this with an explicit payer rail. PM-KISAN: the central government pays ₹6,000/year directly to 11 crore farmers from a ₹60,000 crore budget. MGNREGA: ~₹86,000 crore/year, paid as wages. PMAY-G: subsidies for housing. NHM: ASHA honoraria. Each of these has a ministry, a budget head, a Direct Benefit Transfer rail, an audit. *AI inference for BPL Indians has none of these.*
The IndiaAI Mission’s compute-pricing miracle exists, but, per a *MediaNama* report from April 2026 citing a Rajya Sabha reply by MeitY on February 9, 2026, **only ~₹400 crore of the ₹10,372 crore Mission outlay has actually been released over two years** — ₹21.79 crore in 2024–25 against revised estimates of ₹173 crore, and ₹379.15 crore in 2025–26 against revised estimates of ₹800 crore. Roughly **4% of the five-year corpus, in two years**. The rails are getting built; the budget velocity is glacial; and even within that 4%, almost none has been earmarked for the specific job of *paying for inference consumed by poor citizens*.
The absence of the payer is not a market failure in the textbook sense. It is a *categorical* failure. The inference is currently priced as a B2B sale (Sarvam to Reliance, Sarvam to Tata, Sarvam to a fintech), or as a B2G sale (a foundation-model awardee to a ministry), or as a charity-grant deployment (Digital Green’s Farmer.CHAT funded by Gates, Walmart, and Google.org). What it is *not* yet priced as is a citizen entitlement, the way an MGNREGA day is, or a ration is. Until it is, every individual deployment will hit unit economics it cannot justify and will retreat into grant cycles.
This matters because grant-funded deployments end. ASHABot is one Rajasthan district. Farmer.CHAT — across India, Kenya, Ethiopia, Nigeria — reports about **830,000 users and 6.2 million queries**, real numbers for a grant-funded NGO and very small numbers for India’s BPL population of 234 million. Kisan e-Mitra at **20,000 queries per day** scales to about 7 million queries a year — *one query per BPL Indian every 33 years*. The slope of these projects is real but not on the right vector. Without a population-scale payer, none of them will reach the population.
The natural question is: who *should* pay? The candidates, in declining order of structural fit:
**One**, the Government of India, via a “PM-AI-Sahayak” line item in IndiaAI Mission Phase 2, paying empanelled providers per validated query for BPL Aadhaar holders. ₹1,000–2,000 crore/year buys roughly one query/day per poor household. *This is small money.*
**Two**, sectoral ministries — DA&FW for agriculture, MoHFW for ASHA support, MoRD for MGNREGA navigation — each absorbing inference into their existing schemes. The advantage is ownership; the disadvantage is fragmentation.
**Three**, state governments, the way Rajasthan funded Kisan e-Mitra. Tamil Nadu, Kerala, and Karnataka are credible early adopters; Bihar, UP, and Jharkhand are the actual targets and the laggards.
**Four**, multilateral and philanthropic — Gates, Wadhwani, Walmart, EkStep, J-PAL — bridging to Stage 1 deployments while a public payer is configured. The current default. Inadequate at population scale.
**Five**, the user, via small co-pay (₹5–10/month). Plausible for the upper BPL segments. Implausible for the bottom three deciles.
The cleanest policy move is option one — but India has never, in its post-Independence welfare history, paid per *information transaction*. It has paid for grain, for housing, for school meals, for cash transfers, for hospital admissions. Information has always been a public good produced by the state and consumed by anyone who walks in. AI inference is *the first welfare-relevant information good with a real per-unit cost*. Treating it as a citizen entitlement requires a categorical update — the way Aadhaar required the categorical update of treating identity as something the state issues at unit cost.
I think this update will happen. *I do not think it has happened yet.*
-----
## 6. The error rate is a liability, and the citizen is currently carrying it
A thing I want to mention briefly, because the previous Timlig post on GPU economics was about depreciation, and there is a structurally similar issue here.
Every BPL-facing AI deployment has an error rate. ASHABot, the most carefully evaluated of these, has **doctor-graded accuracy of about 85%** on 163 evaluated responses — meaning roughly 1 in 7 ASHA queries gets a response that a doctor would not endorse. Farmer.CHAT reports a **75% successful-answer rate** in the Singh et al. arXiv paper — meaning 1 in 4 farmer queries goes unanswered or mis-answered. Saaras V3 ASR has a 19.31% WER on the IndicVoices subset and worse on rural dialects — so 1 in 5 spoken words is misrecognized at the front end, before the LLM has even thought.
These are state-of-the-art numbers in the SLM-for-low-resource-languages frontier. They are also unacceptable for a doctor, an agriculture officer, or a scheme grievance system. If the AI tells a pregnant woman that her symptoms are normal when they are actually pre-eclampsia, the household pays the cost. If the AI tells a smallholder that his crop disease is fungal when it is bacterial, the household pays the cost. *The error rate is a real liability, and currently the citizen carries it.*
The depreciation analogy: a commercial neocloud books an H100 over 5 years instead of 3 to make the gross margin look like SaaS, even though the lender amortizes the same chip over 3.5–5 years because that is what the lender actually believes. A BPL-AI deployment books its accuracy as “85%” because that is what the published study says, even though the citizen experiences the misclassification rate as “1 in 7 of my actual real-life questions came back wrong.” The burden of the error rate, in the absence of a recourse mechanism, falls on the user. This is the exact opposite of how welfare-state liability normally works. You cannot refuse food in your ration; you can refuse a wrong AI advisory, but only if you knew it was wrong, which by definition you do not.
The credible architecture is **SLM-augmented frontline worker, not direct-to-citizen autonomous agent**. The ASHA reads the chatbot’s answer, applies clinical judgment, communicates with the patient. The error rate is absorbed by the worker’s training and the hierarchical referral system — exactly the way a junior doctor’s errors are absorbed by a senior. ASHABot’s deployment topology assumes this. Direct-to-citizen agents — Kisan e-Mitra at 20,000 queries/day with no human in the loop — assume the citizen can self-evaluate. For an English-literate user with good context, fine. For a rural Bhojpuri-speaking smallholder asking about cotton bollworm, *not fine*.
This is the second-order constraint that caps how aggressively any of this can be scaled. The unit economics work; the compute is subsidized; the model fits; the language fits. What is not yet built is the **liability and recourse rail** — the AI equivalent of the social audit that MGNREGA has, of the grievance redressal that PM-KISAN has, of the appeals process that ration card disputes have. Without it, scale is reckless.
-----
## 7. What does it actually mean
So what does it mean, if you are a builder, a philanthropist, a state government, or an Anganwadi supervisor, that the price of meaning in Bhojpuri has fallen by an order of magnitude in eighteen months?
A few things.
**One: pick the human cadre, not the citizen.** The unit economics that work today work for SLM-augmented ASHAs, Anganwadi workers, agriculture extension officers, kirana-store-based CSC operators. They do not yet work, with appropriate safety, for direct-to-poorest-citizen autonomous agents. The right question for any deployment is *which existing frontline worker am I making 3× more productive, and how do I measure it*. ASHABot is the right pattern. Scale that, not standalone direct-consumer chatbots.
**Two: assume cloud-mediated voice over WhatsApp/IVR through 2027.** The on-device thesis is real but premature at the median budget Android. Every rupee spent on quantizing a model to run on a Redmi A2 today is a rupee that was not spent on the dialect coverage that would actually serve the user. Optimize for cloud cost (subsidized ₹0.05–0.20 per LLM call), bandwidth efficiency (audio compression, partial offloading of frequent intents), and dialect-tuned ASR. The on-device step is a 2027–28 unlock, not a 2026 one.
**Three: the dialect coverage is the moat.** Sarvam-1 covers ten Indic languages well. Bhojpuri, Maithili, Magahi, Awadhi, Marwari, Santhali, Mundari, Bhili, Gondi, Kui — these are the languages spoken by the actual BPL population, and they remain at fertility and accuracy levels that materially degrade outcomes. Adi Vaani (IIT Delhi consortium with the Ministry of Tribal Affairs, 2025 beta) is the seed. The next ₹100–200 crore of philanthropic money in this space should buy IndicVoices-quality datasets for the eight underserved Indo-Aryan and Munda languages, sourced via Karya-style ethical data labor. *This is the highest-leverage rupee in the entire stack right now.*
**Four: build the payer rail before the application.** The single most under-built piece of this stack is the per-query reimbursement mechanism for BPL citizens. The Aadhaar-authenticated, consent-mediated voucher for inference does not exist; it should. ₹2,000 crore/year, payable to empanelled providers against verified BPL Aadhaar usage, would cover a query a day for every poor household in India. This is one-thirtieth of PM-KISAN. It is well within the IndiaAI Mission’s authorization. *Somebody has to write the GR.*
**Five: instrument outcomes, not adoption.** The next round of BPL-AI evaluation should be J-PAL-quality RCTs comparing SLM-augmented frontline workers against business-as-usual, on hard outcome metrics: institutional delivery rates, ANC4+ completion, agricultural yield per acre, scheme-take-up rates, MGNREGA wage receipt accuracy. The Pratham/J-PAL EdTech evidence base — overall mixed, sometimes negative, often inferior to good human pedagogy — is the relevant reference class. This is the field where good intentions have failed expensively before. The right discipline is to measure outcomes, not query counts.
-----
## 8. What this is, and isn’t
The previous Timlig post on sovereign AI argued that India’s stack is uneven across layers — a direction of travel, sometimes propaganda, often substantive. AI-for-BPL is similar. The slogan — “AI for India’s poor” — papers over a research, distribution, and funding stack with seven distinct layers, of which roughly four are in good shape (model, tokenizer, ASR/TTS, DPI rails), two are in motion (application, distribution), and one is essentially missing (payer).
The honest version is that, *for the first time*, you could give every poor person in India a personal, multilingual, voice-first AI advisor — agriculture, health, scheme navigation — for less than the cost of a single existing welfare program. The model is built. The compute is subsidized. The languages mostly work. The phone, mostly, fits. The frontline workers are in place. What is missing is somebody to write the line item.
That somebody is, almost certainly, the Government of India, because no other actor in this stack has both the rail (Aadhaar, DBT, IndiaAI compute) and the mandate. The fact that this hasn’t happened yet is not a failure of technology or of economics. It is a failure of *categorical imagination* — the same imagination that, twenty years ago, declined to think of identity as a thing the state could provision at unit cost, and ten years ago declined to think of payment rails as a public utility. Both of those happened, eventually, because somebody in Delhi wrote a note.
The two-rupee voice is real. The 234-million-person addressable population is real. The ₹2,000-crore-a-year program that would actually move SDG 1, 2, 3, and 5 needles for a meaningful slice of that population is, today, *not real*. It is a memo away.
-----
## Sources
National Institution for Transforming India (NITI Aayog), *Multidimensional Poverty in India since 2005-06*, Discussion Paper, January 2024 · NITI Aayog, *SDG India Index 2023–24* (4th edition) · UNDP / OPHI, *Global Multidimensional Poverty Index 2024* · Ministry of Statistics and Programme Implementation, Household Consumption Expenditure Survey 2022–23 · *Down To Earth*, RTI reply on the official poverty count (December 2024) · *The South First*, “India has 23.4 crore people living in poverty — highest in the world” · Sarvam AI, Sarvam-1 model card and announcement (October 2024) · Sarvam AI, Saaras V3 ASR and Bulbul TTS API documentation · IndiaAI Mission, Press Information Bureau release on Cabinet approval (PRID 2012355, March 2024) · PIB on IndiaAI compute capacity crossing 34,000 GPUs (PRID 2132817) · *MediaNama*, “IndiaAI Mission: Only Rs 400 Crore Released in Two Years” (April 2026), citing Rajya Sabha reply by MeitY (February 9, 2026) · SME Futures, “Rs 65 per GPU per hour: Subsidy rate under India AI Mission from 14 service providers” · Outlook Business, “Over 17,000 GPUs successfully installed Under Govt’s IndiaAI Mission” · AI4Bharat (IIT Madras): IndicTrans2 (arXiv 2305.16307), IndicVoices (arXiv 2403.01926), Airavata, IndicConformer, “Bhojpuri and Hindi Rural Women ASR” Hugging Face collection · OpenAI, GPT-4o tokenizer announcement; Microsoft Azure technical commentary on `o200k_base` for Indic languages · Microsoft Research / Khushi Baby, ASHABot evaluation, CHI 2025 (Ramjee et al., “ASHABot: An LLM-Powered Chatbot to Support the Informational Needs of Community Health Workers”) · Digital Green, Farmer.CHAT public deployment metrics; Singh et al., “Farmer.Chat: Scaling AI-Powered Agricultural Services for Smallholder Farmers” (arXiv 2409.08916) · Wadhwani AI, CottonAce program data and Google.org / Welspun / Deshpande Foundation assessments · BBC Media Action / ARMMAN / MoHFW, Kilkari maternal mHealth program · Verma et al., “Leveraging AI to improve health information access in the World’s largest maternal mobile health program,” AI Magazine (Wiley 2024) · Microsoft Research / EkStep / AI4Bharat / OpenNyAI / Bhashini, Jugalbandi · PIB, “KISAN E-MITRA and IoT enabled systems to improve crop productivity” (PRID 2117392) · IndiaAI, “Exploring Pradhan Mantri-KISAN AI Chatbot” · Ministry of Rural Development, MGNREGA wage notification FY 2025–26 (effective April 1, 2025) · Ministry of Health & Family Welfare, ASHA honorarium notification (Lok Sabha unstarred Q4764, March 28, 2025) · Bhashini documentation; PIB, “BHASHINI: Transforming Maha Kumbh through Multilingual Innovation” (PRID 2093333) · BharatNet, 2.15 lakh gram panchayats coverage (Department of Telecommunications, December 2025) · ASER Centre, *Annual Status of Education Report 2024* · National Family Health Survey-5 (2019–21) · *Data For India* on phone access and internet penetration · Ookla, India connectivity report H1 2025 · J-PAL, EdTech Sector Review · Banerjee et al., Mindspark CAL evaluation · Pratham, *Teaching at the Right Level* program reports · Karya / DRK Foundation public profiles · Rajya Sabha, IndiaAI Mission foundation-model awardee replies (MoS Jitin Prasada, February 13, 2026) · MeitY, *India AI Governance Guidelines* (November 2025) · Press releases from Sarvam AI, BharatGen consortium, Hanooman / SML / BharatGPT, Krutrim, Soket AI, Gnani AI, Gan AI · Adi Vaani consortium press materials (Ministry of Tribal Affairs, 2025) · Tribal language demographics, 2011 Census · arXiv 2512.06490, “Optimizing LLMs Using Quantization For Mobile Execution” · arXiv 2410.03613, “Large Language Model Performance Benchmarking on Mobile Platforms” · arXiv 2506.09653, “Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women” · IBM Granite, Microsoft Phi-3, Google Gemma, Meta Llama 3.2 documentation · *Local AI Master*, on-device SLM benchmarks 2026 · Kettani & Moulin, “Rethinking the Role of Technology for Development in the AI Era: From AI4D to Smart ICT4D” (IJETT 2025) · The George Institute for Global Health, “AI for Community Health Workers in India” series · Frontiers in Global Women’s Health, on Kilkari deployment in Assam (2025) · Inc42, *From LLMs to Verticalisation: India’s Sovereign AI Models Take Shape*.
---
# Sovereign AI in India: A Word in Search of a Definition
Source: https://www.timlig.com/posts/sovereign-ai-india/
Published: 2026-05-04
Tags: AI, India, sovereign AI, policy, semiconductors, infrastructure
## I. The Word
Here is a fun thing about “sovereign AI” in India: nobody knows what it means.
At a major AI summit in Delhi in February 2026, the word was attached to: a GPU rental service, a cloud product from a large American software company, a 105-billion-parameter language model trained on chips imported from Taiwan via California, a state-government industrial park, and the entire stack of Indian digital public infrastructure that has been around since well before anyone was calling things sovereign. A startup booth advertised “instant, sovereign GPU access.” The GPUs were made in Taiwan. A large American software company offered a “Sovereign Cloud,” which is the same software they sell everywhere else, with a sticker on it. The 105-billion-parameter model is, I should say, a very good model. It is also sovereign.
The thing you would normally do here is look up the legal definition. There isn’t one. The Information Technology Act does not define sovereign AI. The 2023 data protection act does not define sovereign AI. The Cabinet note that approved a roughly $1.25 billion mission for AI in March 2024 mentions “tech sovereignty” the way a wedding invitation mentions the weather. The 2018 national AI strategy paper does not use the term at all; it calls India the “AI Garage for 40% of the world.” A garage, you’ll notice, is where someone else parks their car.
So the situation is: a word everyone is using, attached to a great deal of capital, with no operational definition. To the IT ministry, sovereign AI means national capability. To a cloud provider, it means a data center inside Indian borders. To a chip startup, it means RISC-V. To a foundation-model company, it means tokens that handle Tamil better than GPT-4. To a procurement officer, it means a category they can buy from. To a venture capitalist, it means a category they can sell into. Everyone agrees sovereign AI is a good thing, while disagreeing about what it is. The agreement holds because of the disagreement.
The closest thing to an authoritative definition came, fittingly, from a chip company. The CEO of a large American GPU manufacturer gave a speech in Dubai in early 2024 saying every country needs to “own the production of its own intelligence” and helpfully suggested codifying one’s national language and culture into a large language model. This is, if you squint, a perfectly reasonable thing for a chip CEO to say, in the same way it would be perfectly reasonable for the CEO of a cement company to argue that every country needs to own the production of its own roads. It is, conveniently, a definition under which sovereign AI is sold by the speaker’s company. It has been quietly adopted, in various forms, by basically every AI policy document worldwide that does not have its own definition. India’s documents are among them.
The fuzziness is not a bug. It is the strategy. As long as everyone can read their preferred meaning into the word, everyone is on the team. What I want to do in this piece is walk down the AI stack, layer by layer, and ask where India is actually building something, where it isn’t, and where the money is going. The answer is uneven, but in a more interesting way than “uneven” suggests.
## II. The Stack
It is convenient to think of AI as a stack. The lowest layer is electrons. The highest layer is whatever the AI is doing for you. We will go bottom to top.
### Energy
Electricity is not usually part of the AI sovereignty conversation, which is strange because every other layer assumes it. India’s installed capacity crossed 500 GW at the end of 2025, with renewables around 190 GW. Building a data center in India costs about $7 per watt — against $10 in the United States, $11 in the United Kingdom, and roughly $6 in China. India is the second-cheapest large economy in the world to build AI compute capacity in. Nobody at the summit was selling sovereign electrons, but every gigawatt-scale AI campus that gets announced is, on inspection, a bet on the cheap-power thesis. The grid is uneven across states, AI workloads want 99.999% uptime, and storage is thin, but as foundations go, “we have lots of cheap electricity” is a pretty good one.
### Silicon
Here is where the rhetoric and the reality have the most fun together.
India has approved more than ten semiconductor projects since 2021, with cumulative committed investment north of $18 billion. The flagship is a roughly $11 billion fab in Gujarat, built by a domestic conglomerate with a Taiwanese partner, planned at 50,000 wafer starts per month. There is a $3.2 billion assembly-and-test facility in Assam, smaller units from an American memory company, two Indian conglomerates, a Taiwanese contract manufacturer, and an Israeli foundry. First chips from the flagship are targeted for late 2025; stable production by 2026.
The chips will not be AI chips. Every approved Indian fab is at *mature nodes* — 28 nm and above. Frontier AI silicon is fabricated at 5, 4, and now 3 nanometers, at facilities in Taiwan and South Korea, on equipment from a Dutch monopolist that the United States has restricted from selling to China and that nobody has tried to sell to India because India has not asked. The gap is roughly five process generations. Nobody seriously expects India to close it this decade.
This sounds like a problem and is mostly not. The fabs being built are the right fabs — they will supply the chips that go into cars, appliances, telecom equipment, defence, and the enormous ecosystem of devices that does not need bleeding-edge silicon. They are not, in any meaningful sense, AI fabs, and the official position is admirably honest about this. Any Indian-designed AI accelerator, including the indigenous GPU programme announced under the AI mission with a 2029 production target, will be fabricated abroad. The mission is buying option value at this layer, not catching up.
### Chips and Accelerators
One layer up is chip design, where India has been doing real work for a while.
The open-source RISC-V instruction set, for reasons that are partly technical and partly geopolitical, has become the default architecture for Indian processor design. Two academic-led cores anchor the ecosystem; both have had multiple successful tape-outs at older nodes. A clutch of fabless startups now constitutes the credible Indian footprint in AI hardware: a microcontroller company that shipped the first commercially designed Indian chip in May 2024 ($8 million Series A), a RISC-V core company spun out of the same institute, a gallium nitride defence-and-telecom company in Bangalore, a neuromorphic accelerator company building reconfigurable AI hardware, and a consumer-internet founder who has committed roughly $230 million of family-office capital to an AI venture planning to tape out an indigenous AI chip by 2026.
None of these companies is competing with the frontier. The frontier in 2026 is roughly 80-billion-transistor accelerators with millions of high-bandwidth memory stacks attached, on a process that one company in Taiwan can run. Indian startups are building the tier below — application-specific accelerators, edge AI inference, low-power IoT controllers, vision processors. Taiwan, you will recall, started in the lower tiers and climbed up. India has the option. The question is whether it takes the climb seriously.
In the meantime, every GPU in every Indian AI training cluster — including every GPU in the much-celebrated subsidised national compute portal — is American. There are roughly 38,000 of them, supplied primarily by one company in Santa Clara and secondarily by a competitor in Sunnyvale. They are foreign silicon, fabricated in Taiwan, running CUDA software that is also American. Calling this stack sovereign requires a generous definition of sovereign — basically the definition under which “buying things” is a form of “owning things” — but the *deployment* is sovereign, in that India decides who gets to use the GPUs and at what price. This is a meaningful kind of sovereignty. It is not the kind the word evokes.
### Compute Infrastructure
This is the layer where the headline numbers look best.
The national AI compute portal, approved in March 2024, originally targeted 10,000 GPUs. By late 2025 it had onboarded roughly 38,000 across fourteen empanelled providers. Pricing is the striking part. About ₹65 per GPU-hour on average, after a 40% government subsidy. A startup or researcher can rent an H100 in India for under $1.50 an hour, several multiples below global commercial benchmarks. The Indian government has effectively created the cheapest production AI compute environment in the world for domestic users by combining cheap electricity, public-private financing, and a procurement subsidy. A country with a small AI public budget cannot finance frontier model training. It *can* make frontier compute affordable to people who would otherwise be locked out of it. India did the second thing.
The data centre buildout is, separately, dramatic. Commercial capacity stood at about 950 MW at the end of 2024 and is projected to roughly double by the end of 2026 and reach 9 GW by 2030. A domestic energy and telecom conglomerate has announced a 1 GW campus in Gujarat, scalable to 3 GW. An American search company in partnership with an Indian infrastructure conglomerate has announced a $15 billion gigawatt-scale build on the eastern coast. The three large American hyperscalers have together committed more than $50 billion through 2030. The Union Budget for 2025–26 extended the data-centre tax holiday to 2047. Tax-advantaged power-hungry capital infrastructure with 22-year fiscal certainty is, it turns out, a terrific business.
There is, however, a question the marketing materials skirt: when an American hyperscaler operating in India receives a legal demand from its home government for data stored on Indian soil, what happens? The hyperscaler will say it complies with both jurisdictions. The jurisdictions may give conflicting instructions. The “sovereign” suffix on the cloud product is a self-certification with no Indian regulatory standard to certify against. This is unresolved, and it is also unresolved in the EU under the CLOUD Act, where they have been working on it for years. You cannot regulate a layer that does not exist yet. First the data centres, then the doctrine.
### Networking
India is structurally well-placed here. 5G commercial rollout took roughly two years and produced the world’s second-largest subscriber base. The rural fibre programme has connected hundreds of thousands of villages. Submarine cable landing stations are being added. The 2023 data protection act establishes a soft data-localisation regime with substantial extraterritorial reach. None of this is glamorous. All of it works.
### Data
Here is where India is genuinely distinctive.
Three asset classes matter. The official datasets platform launched under the AI mission now hosts more than 5,500 datasets across 20+ sectors and has a five-year allocation of about ₹200 crore. The national language platform hosts more than 350 AI models, has had over a million downloads, and has signed institutional MoUs with the Indian railway system for voice-based translation. The most important asset, for my money, is the academic-led open language ecosystem from a southern technical institute, which has produced a 251-billion-token pretraining corpus across 22 languages, a 74.7-million-pair instruction-tuning dataset, and multilingual TTS datasets covering all 22 constitutionally recognised Indian languages. This is real data, generously licensed, in languages that are otherwise badly underserved by global AI.
Here is the wonky part, which I find genuinely cool. Global frontier models are bad at Indian languages — not because they cannot learn them, but because their tokenizers were trained on English-heavy corpora and consequently chop Indic-language inputs into 4 to 8 tokens per word, against 1.4 for English. Inference in Hindi or Tamil or Bengali is therefore roughly five times more expensive than inference in English on the same model. People have started calling this “the token tax.” It is the single most concrete economic argument for an Indic-first foundation model: the moat is not parameter count, it is tokenizer design plus high-quality language corpora plus inference pricing. India has all three.
The 2023 data protection act is awkwardly aligned with all this — it is consent-centric, written for individual transactional data, not for trillion-token web scrapes. The AI industry’s lobbying body has formally asked for a research and training exemption for “publicly available data.” The request is pending. Whether the act gets amended, a separate framework gets written, or the data fiduciaries figure out a workable interpretation — these are the eighteen-month questions to watch.
### Foundation Models
This is where the political bet is concentrated.
The AI mission’s foundation-models pillar received more than 500 proposals in its first year. Twelve teams have been selected across two phases, with allocations ranging from a few hundred thousand dollars to about $125 million for the largest awardee, an academic-led consortium based at a Mumbai technical institute. The largest commercial winner received about ₹247 crore for a 120-billion-parameter open-source model branded as India’s “sovereign LLM ecosystem.” Other awardees cover speech recognition, healthcare reasoning, generative video, and Indic translation.
Models actually shipped at the February 2026 summit included a 30-billion and 105-billion-parameter mixture-of-experts pair from the largest commercial awardee (32K and 128K context windows respectively), a 17-billion-parameter multimodal model from the academic consortium supporting all 22 official languages, a low-latency speech-to-text and text-to-speech pair with claimed character error rates below 0.6%, and a verticalised health-reasoning model. Outside the AI mission, a consumer-internet founder’s AI venture released two models in late 2024 and early 2025 — a 7-billion and a 12-billion-parameter model — and has committed roughly ₹10,000 crore over the year to its AI program.
None of these models is competitive with the global frontier on parameter count, training compute, or general reasoning benchmarks. The gap is widening, as frontier labs train on tens of thousands of GPUs for months at a time with budgets larger than the entire AI mission. If the goal is to win every leaderboard against GPT-5 and Claude 4 and Gemini Ultra, India will lose, by an enormous margin, and so will every other country, including most American countries, which are also America.
But that is not the goal. The official thesis is that foundation models are becoming commodities, that 50-billion-parameter models will handle 95% of Indian use cases, and that the strategic value is in Indic-language reasoning, voice-first interfaces, and domain-specific deployments rather than chasing the frontier. This is plausible, and it is also, conveniently, the bet India can afford. If it works, it is *better* than chasing the frontier, because chasing the frontier is a game where the second-place finisher gets nothing. Specialisation is a more durable competitive position than scale at any given price point. The entire history of computing tells you this. The Indian bet is on specialisation.
A separate thing worth flagging. The largest commercial awardee disclosed in April 2025 that a central government body would take an equity stake in the company in exchange for compute resources. By global AI funding standards this is unusual, and other AI founders have asked, in essence, why this firm and not theirs. The answer is presumably that the firm in question had the right team and architecture at the right time, but it would help — and the policy community has been gently suggesting this — for future rounds to publish selection criteria up front. This is a normal piece of process improvement.
### Applications
If you only read one section of this article, read this one.
India has built, over the last decade, a population-scale digital public infrastructure. Identity (a biometric system covering 1.3 billion people). Payments (a real-time rail running tens of billions of transactions a month). Documents (a government cloud document wallet). Commerce (an open digital commerce protocol). Languages (the multilingual platform mentioned above). Account aggregation (a consent-based financial data sharing framework). None of this was built for AI. All of it is now AI-relevant, in the way that the Roman road network turned out to be Christianity-relevant several centuries after it was built. A health-AI startup deploying in India does not have to build identity, payments, or document verification — it plugs into existing rails. The friction cost of deploying AI to a billion people is, in India, dramatically lower than in any other country in the world. This is enormous, and it is permanently true.
The AI mission’s application pillar (₹689 crore) has selected 30 applications across agriculture, health, climate, and disaster response. Sectoral hackathons have been run with the cyber crime coordination centre, the geological survey, the alternative medicine ministry, the small-enterprise ministry, the financial reporting authority, and the national cancer grid. There is an AI Centre of Excellence for Education with a ₹500 crore allocation in the 2025–26 budget. The defence ministry is integrating Indian-designed RISC-V chips into satellite and avionics systems for fault tolerance. None of this is AI sovereignty in the chip sense. All of it is AI sovereignty in the more useful sense — a country building things, on its own rails, in its own languages, for its own users, faster than any other country at this scale could. This is the third path between American private platforms and Chinese state platforms, and the rest of the Global South is watching.
### Governance
India has explicitly chosen not to enact a horizontal AI Act of the European kind, and this is a defensible policy choice rather than a missing piece of homework.
The architecture is four-pillared. Existing horizontal law applied to AI (the IT Act, the criminal code, the consumer protection act, the copyright act, the data protection act). Ministry-level advisories — the most famous of which was issued on March 1, 2024 and rescinded on March 15, 2024 after the AI startup community made its views known. The November 2025 Governance Guidelines, organised around seven principles, principle-based and explicitly non-statutory, with an AI Governance Group, a Technology and Policy Expert Committee, and an AI Safety Institute. And sectoral regulators retaining domain oversight — the central bank for finance, the markets regulator for securities, the medical research council for health.
The case for this architecture is straightforward: AI is moving too fast for primary legislation, sectoral regulators understand their sectors, and a principles-based framework leaves room to adapt. The case against is that voluntary guidelines have no teeth and that something will eventually go wrong. The European Union has the world’s most comprehensive AI Act and also, by most measures, the world’s least competitive AI industry. Causation is hard to establish, but the correlation should give pause to anyone who thinks “more rules” is the obvious answer.
India hosted the 2026 summit, renamed from “AI Safety Summit” to “AI Impact Summit,” which is a substantive choice — the renaming reflects the global shift away from existential-risk framing toward implementation and deployment. The Delhi Declaration is organised around People, Planet, and Progress. It is the first major AI governance document from the Global South, and it signals a real reorientation of the international conversation.
## III. The Money
Let’s talk about the budget for a minute, because it tells you what the strategy actually is.
The five-year mission allocation, approved in March 2024, is about ₹10,372 crore — call it $1.25 billion. It is divided across seven pillars: 44% for compute capacity, 19% for foundation models, 19% for startup financing, 9% for skills, 7% for application development, 2% for the datasets platform, 1% for overheads, and 0.2% for safe and trusted AI. The safe-and-trusted-AI line item is one-fifth the size of the overheads-and-contingency line item. The signal is: India is funding capability, not governance, and is doing so on purpose. You may agree or disagree. It is at least clear.
As of an analysis published in April 2026, roughly ₹400 crore had actually been released — about 4% of the five-year outlay, in two years. This can be read two ways: either the mission is underspending, which is a problem, or the mission is ramping into capability rather than dumping money into a market that has not yet absorbed it, which is good fiscal hygiene. I lean toward the second. The next phases require talent, institutions, and deployment, none of which are accelerated by spending faster than the ecosystem can use.
Meanwhile, the three large American hyperscalers have together announced more than $50 billion of Indian cloud and AI investment over the next several years. That is roughly forty times the central mission’s full five-year corpus. Public investment in the sovereign-AI agenda is dwarfed by foreign private investment in Indian compute infrastructure. This is the comparison that gets made to suggest the mission is undersized, but the mission was never going to fund the build-out of the compute layer. That was always going to be private capital. The mission was always going to be a coordinator and a subsidiser.
What the mission *is* funding directly is the layers where private capital won’t go: public datasets, Indic-language model training, application-layer hackathons in the alternative medicine ministry, skilling programs at tier-2 polytechnics, the 27 IndiaAI Data and AI Labs in tier-2/3 cities. None of these is a margin business. None of these gets built by AWS or Microsoft or Google. The mission is doing exactly what an industrial policy mission is supposed to do — funding the public goods that the private market won’t, and letting the private market handle the layers where it has the comparative advantage. Dressed up in maximalist sovereignty rhetoric, this is, at its core, remarkably orthodox industrial policy.
## IV. So What Does It Mean
Sovereign AI in India is not a definition. It is a direction of travel. The travel is uneven because different layers of the stack have different physics. India will not be making 3-nanometer logic this decade, and pretending otherwise would be silly, and nobody important is pretending otherwise. India *is* building the data, the languages, the digital public infrastructure, the deployment rails, and an increasingly capable indigenous foundation-model and chip-design ecosystem on top.
The honest framing is that full-stack AI sovereignty is structurally infeasible for any country other than perhaps the United States and China, and the realistic objective for India is *layered* sovereignty: own the layers you can credibly own, partner on the layers where partnership is strategic, accept dependence where dependence is honest, and preserve switchability across providers and jurisdictions. This is not autarky. It is not the Chinese model, which is not replicable anyway. It is something genuinely new, and it has the unusual property of being *better* than the alternatives because India has DPI and 22 official languages and 1.4 billion people and cheap power and a deep talent pool, and none of those things are accidents of strategy. They are the inputs.
The slogan is sovereign. The stack is interdependent. The capital is mostly private, mostly foreign, mostly in compute. The genuinely sovereign assets — the languages, the data, the digital public infrastructure, the application-layer policy choices — are the ones built over decades, often by people who were not using the word sovereign when they built them. The AI mission is now layering on top.
If you ask what sovereign AI in India means in 2026, you get a different answer from every person you ask. If you ask what it will look like in 2030 — what will exist that does not exist today — most of the answers converge. There will be Indic-language models running on Indian compute, embedded in Indian DPI, used by Indian citizens, with broadly Indian governance, at prices no other country can match. You can call it sovereign if you like. You can call it whatever you like. The country is going to build it either way.
-----
## Sources
- NVIDIA blog and corporate materials, including the January 2024 World Governments Summit address in Dubai.
- Press Information Bureau release on the Cabinet approval of the IndiaAI Mission, March 7, 2024 (PRID 2012355).
- Press Information Bureau release on IndiaAI Compute Capacity crossing 34,000 GPUs (PRID 2132817).
- Press Information Bureau release on the Tata Electronics–ISM Fiscal Support Agreement, March 5, 2025 (PRID 2108602).
- NITI Aayog, *National Strategy for Artificial Intelligence* (2018).
- NITI Aayog, *Responsible AI for All* discussion papers (2021–22).
- MeitY, *India AI Governance Guidelines* (November 2025).
- Office of the Principal Scientific Adviser, *Democratising Access to AI Infrastructure* (white paper, 2025).
- RBI FREE-AI Committee Report (August 2025).
- MeitY AI Advisories of March 1, 2024 and March 15, 2024.
- Draft IT Intermediary Rules amendments on synthetic content and deepfakes (October 22, 2025).
- Digital Personal Data Protection Act, 2023, and DPDP Rules notified in 2025.
- Union Budget documents, 2024–25, 2025–26, and 2026–27.
- Observer Research Foundation, *Operationalising India’s Sovereign AI Stack: From Intent to Capability* (2025).
- Carnegie India / Carnegie Endowment for International Peace, work on India’s semiconductor mission and AI compute strategy (2024–2025).
- Takshashila Institution, *Building India’s Data Centres* (October 2025).
- Tony Blair Institute for Global Change, *Sovereignty in the Age of AI: Strategic Choices, Structural Dependencies and the Long Game Ahead*.
- Brookings Institution, *Sovereignty, Safety, and Scale: Takeaways from the India AI Impact Summit* (2026).
- Chatham House, *How Middle Powers Can Weather US and Chinese AI Dominance* (February 2026).
- European Commission documents on the AI Continent Action Plan and Strategic Compass.
- EY India, *The AIdea of India 2026: Sovereign AI in India*.
- The Ken, *India Called Its AI Sovereign. The US Government Can Still Access It*.
- MediaNama, *IndiaAI Mission: Only Rs 400 Crore Released in Two Years* (April 2026).
- Rest of World, *The Myth of Sovereign AI: Countries Rely on US and Chinese Tech*.
- Avasant, *The Illusion of AI Sovereignty: Washington and Beijing Still Pull the Strings*.
- Lawfare, *Sovereign AI in a Hybrid World: National Strategies and Policy Responses*.
- Inc42, *From LLMs to Verticalisation: India’s Sovereign AI Models Take Shape*.
- AI4Bharat technical reports and dataset releases (IIT Madras), including IndicBERT, IndicBART, Airavata, Bhasha-Abhijnaanam, Rasa, and Setu.
- Bhashini platform documentation and partnership MoUs.
- AIKosh / IndiaAI Datasets Platform documentation.
- Public materials from foundation-model awardees and AI ventures.
- C-DAC documentation on the National Supercomputing Mission, AIRAWAT-PSAI, and PARAM Rudra systems; Top500 list, November 2025 edition.
- Press releases and funding announcements from Mindgrove Technologies, InCore Semiconductors, AGNIT Semiconductors, and Morphing Machines.
- Future of Privacy Forum, *Five Ways in Which the DPDPA Could Shape the Development of AI in India*.
- Internet Freedom Foundation, *Analysis of the 2026–2027 Budget*.
- India AI Impact Summit 2026 reportage and the text of the Delhi Declaration.
- News reports from Business Standard, Business Today, The Tribune, TechCrunch, Storyboard18, Dholera Times, Indian Masterminds, and Trade Brains.
---
# The Cheapest GPUs in the World
Source: https://www.timlig.com/posts/cheapest-gpus-in-the-world/
Published: 2026-04-27
There is a fact about Indian AI infrastructure that, if you sit with it for a minute, is genuinely funny.
The Indian government, through its IndiaAI Mission, will currently sell you an hour of an H100 GPU for about **₹65**. That is roughly **78 cents**. If you happen to be a startup building an "indigenous foundational model," the government will sell you that same H100 for **zero rupees**, because the compute subsidy in that case is 100%. *One hundred percent.* You bring the engineers, the state brings the silicon, the silicon costs you nothing.
Meanwhile, if you walked into a commercial Indian neocloud and asked for the same H100 SXM5 on-demand, the sticker price would be around **₹249/hr**, or about $2.99. If you walked into the AWS Mumbai region and asked the same question, you'd be quoted something north of ₹330/hr.
So the price of one (1) H100-hour in India, depending on who you ask, is somewhere between **zero and four dollars**. This is a 4× spread on what is supposed to be the most fungible commodity in the entire AI stack. *The same chip. The same hour. Different invoice.*
The natural question is: which one of these prices is the *real* price? And the answer, which I want to spend the next 3,000 words on, is that none of them are. The real price is something else entirely, and almost nobody in India is currently set up to measure it, because we have collectively decided that GPUs are like roads — public goods you pour money into so that the *next* layer of the stack can do something interesting — rather than like, you know, **business assets that are supposed to make money**.
This is fine! It might even be smart industrial policy. But it does mean that when you read about the Indian AI ecosystem and someone tells you their AI startup has a 70% gross margin, you should know that you are reading a sentence with approximately the same epistemic content as the sentence "my house is worth a lot because I really like it."
Let me explain.
---
## 1. The basic problem
The basic problem is that an AI company's gross margin is almost entirely a function of two things you cannot see on the income statement.
The first is the **utilization rate** of the underlying GPU — what fraction of the time the chip is actually doing math that someone is paying for, as opposed to (a) sitting on, (b) reading and writing memory while the tensor cores are idle, (c) waiting for the next batch, (d) checkpointing because GPU #14,332 in the cluster just died, or (e) crunching numbers for an internal experiment that will be deleted in six weeks.
The second is the **depreciation schedule** — how many years you, the operator, have decided this $40,000 chip will keep earning revenue. If you say "two years," you have to recognize $20,000 of expense per year and your gross margin looks bad. If you say "six years," you only recognize $6,667 per year and your gross margin looks like SaaS. Same chip. Same revenue. Different number on the page. Investors trade these companies at different multiples based on the number on the page.
This is true everywhere. It's true in the United States, where a leading hyperscaler quietly extended GPU useful life from 4 to 6 years between 2022 and 2024 and added something on the order of $3 billion to annual operating income — purely from the accounting change, not from any chip getting better at its job. It's true in Europe.
But it is *especially* true in India, because in India there is a third variable that the US and Europe don't have, which is **the government writing checks for the GPU bill**. And this turns out to do strange things to the math.
---
## 2. The IndiaAI subsidy and the breakeven that doesn't exist
Here are the numbers, very quickly.
IndiaAI Mission's total budget is **₹10,371.92 crore** — about **$1.14 billion** — over five years. Of that, **₹4,563 crore** (~$500M) is earmarked for compute. As of February 2026, the Mission has onboarded **38,000+ GPUs** across 14 empanelled providers, with another 20,000 in the pipeline. The lowest accepted bid in the most recent tender was **₹65/GPU-hour** as a baseline rate; H100s specifically came in around **₹92/hr**. The government then layers on a **40% subsidy** for general approved users and a **100% subsidy** for a select group of startups developing foundational models. The Minister has called it "the cheapest compute facility in the world," which is the kind of thing politicians say but which, in this case, is approximately *true*.
It is also, to a first approximation, **free money for compute**. And free money for compute does what free money always does: it shifts the breakeven analysis from *"at what utilization does this GPU pay for itself"* to *"how much of this should I grab before they notice."*
Let me be more precise. A leveraged commercial Indian neocloud — no subsidy, debt-financed, importing the chip with 18% IGST stacked on top (creditable for GST registrants but a working capital drag) — needs roughly **75–90% utilization at $2.20/hr realized rates** to break even, depending on whether you depreciate the H100 over 4 or 6 years. Skinny margin, fragile to power outages, fragile to a single customer leaving, fragile to NVIDIA shipping Blackwell at scale.
A subsidized IndiaAI provider, by contrast, has a "breakeven" that is essentially **whatever rate the government has agreed to pay**, plus whatever they can sell on the side. The economics are not "is this asset productive" but "did I win the tender." Which is fine for a public-goods program, except for one small detail, which is that **nobody — not the government, not the providers, not the startups burning the free hours — has any structural incentive to measure goodput**. The chips are paid for. The hours are paid for. *Why would you care if your MFU is 22%?*
The cleanest illustration of this incentive structure is the fact that, when the IndiaAI tender for 2,400 H100s went out, the **AWS Managed Service Providers declined to match the lowest bid**. This is in some ways the most interesting data point in the entire program. AWS has these chips. AWS would presumably like the revenue. But AWS knew that participating at the floor price would set a reference rate that would haunt the rest of its commercial book in India forever, and they walked away. The dominant US hyperscaler has decided that the marginal IndiaAI tender is worth less to it than not signaling that an H100-hour can ever cost less than $2. Which is worth thinking about, because it tells you what they think the *real* price is.
---
## 3. The thing nobody is measuring
Okay. So now, having established that nobody in the Indian AI ecosystem has a strong economic reason to measure GPU usefulness, let's talk about what that usefulness actually looks like when someone bothers to measure it.
The single most important number in this entire essay is **38–43%**.
That is the **Model FLOPs Utilization** achieved by one of the world's most sophisticated AI infrastructure teams on a 16,384-H100 training cluster, running a flagship 405-billion-parameter model, over a 54-day continuous training window. It is one of the highest publicly disclosed MFU figures for a frontier-scale run. It is the *good* number.
To translate: the chips were doing the floating-point operations the model needed for **about 38–43%** of the wall-clock time the customer was paying for. The other 57–62% was spent reading memory, waiting on communication, recovering from failures, and doing accounting work that doesn't show up in the loss curve. This is not a bug. This is the **state of the art**. Most production inference deployments run at 20–40% utilization. Most enterprise fine-tuning workloads, when measured honestly, run lower than that.
There is a separate, related number that I find delightful, which is the gap between what `nvidia-smi` reports and what is actually happening. **You can read 100% GPU utilization from `nvidia-smi` while doing zero floating-point operations**, because the metric counts memory reads as "utilization." A consulting engagement at one foundation model company found a workload showing 100% nvidia-smi utilization and 20% MFU. After optimization: 38% MFU, 4× wall-clock speedup, *same dashboard reading.* The dashboard was lying. The dashboard is always lying.
And then there's the failure rate. The same 16,384-GPU training run mentioned above logged **419 unexpected interruptions in 54 days** — one every three hours — and **58.7% of those interruptions were GPU-related**. The CPUs failed twice in 54 days. The $30,000 accelerators failed 246 times. The cheap part is reliable. The expensive part is the unreliable part. *This is the thing you bought.*
For Indian operators, this matters more than it does in the US, for a reason that is genuinely structural and that I want to dwell on.
---
## 4. The tropical PUE tax
Power Usage Effectiveness — PUE — is the ratio of total energy a data center consumes to the energy that actually goes into the compute. A PUE of 1.0 means every joule of electricity is being turned into computation. A PUE of 2.0 means half your power bill is air conditioning.
Here are some PUE numbers, roughly:
- **Nordic hyperscaler data center (Finland):** ~1.10
- **US hyperscaler data center (Texas/Iowa):** ~1.15–1.20
- **Frankfurt commercial colocation:** ~1.30–1.35
- **Mumbai industrial data center:** **1.55–1.70**
- **Bengaluru/Chennai (renewable-optimized):** ~1.45
What this means in practice is that for every 700-watt H100 you run in Mumbai, you pay for **1,085 watts** of grid power, where the same chip in Stockholm runs on 770 watts. The Indian "cheap power" story — Maharashtra HT industrial tariff at around ₹8.36/kWh, about $0.10/kWh, genuinely cheaper than Frankfurt — gets *almost entirely consumed* by the cooling penalty. Run the math: cheaper electrons, more electrons needed, net result roughly a wash with US Texas and meaningfully *worse* than Nordic.
This is the part that I think gets underappreciated in Indian AI infrastructure conversations. The bull case for India is: cheap power, cheap labor, sovereign demand. The bear case is: the cheap-power claim is half-true at best because the cooling overhead eats it, the cheap-labor claim works for FinOps and engineering but not for the chip itself (you are paying world prices, in dollars, after an 18% IGST detour), and the sovereign-demand claim works only as long as the government keeps writing the checks. *Which it might. But it might not.*
The single biggest cost lever an Indian operator actually has — the one that is genuinely structural and not vulnerable to subsidy reform — is **labor**. A senior ML platform engineer in Bengaluru costs, on the high end, about $52,000 a year. The equivalent role at a US frontier lab is somewhere between $300,000 and $800,000 in total comp. A five-person Indian FinOps team costs ~$200,000 fully loaded. The American equivalent is ten times that.
Which leads to a question: if FinOps headcount is essentially free in India relative to compute, why does almost no Indian AI company run a serious goodput-measurement program? And the answer, again, is that nobody is asking them to. The hours are free. The chips are subsidized. The customers don't know the difference between MFU 20% and MFU 40%, and nobody on the cap table is making them care. *Yet.*
---
## 5. The IT services time bomb
While we're here, I want to talk about the single biggest distortion in the Indian AI economic story, which is the **IT services sector**.
Top-five Indian IT services firms collectively lost **more than $150 billion in market capitalization in the first nine months of 2025**. One major firm announced its first-ever mass layoff in July 2025: **12,000 jobs**. ICICI Direct estimates AI may cause **2–3% annual deflation in traditional IT services revenue** going forward. The traditional Indian IT model is the world's largest, longest-running labor arbitrage trade — Indian salaries, US bill rates, take the spread for forty years — and Copilot is currently eating the spread.
And here is the cruel irony: the Indian IT services sector is *built* to do FinOps. It has the people, the processes, the SLA-discipline DNA, the fluency in client billing structures. Indian GCCs of foreign multinationals genuinely lead the world in some of this. The skill is there. But the IT services *companies themselves* have the wrong economic structure to capture AI infrastructure margin — they sell hours, not outcomes, and the hours are getting cheaper because their tooling is getting better at writing code, which is the thing the hours used to do.
The companies that *can* capture the margin — the AI-native Indian startups — are the ones currently being subsidized to *not* worry about margin. Which is a very strange equilibrium.
If I had to pick a single sentence to summarize the strategic situation of Indian AI infrastructure in 2026, it would be: *the entity best equipped to discover the real cost of AI is being killed by AI, and the entity in the best position to ignore the real cost of AI is being kept alive by the government on the explicit condition that it ignore the real cost of AI.*
---
## 6. The depreciation schedule, briefly
I should at least mention the depreciation thing, because it's the part of this story that is going to matter when, eventually, the Mission scales back, or the subsidy expires, or the government runs another tender at a tighter price, and somebody actually has to write down the asset.
H100 cards in India today are being booked across operators at **5–6 year useful lives**, mostly mirroring US neocloud and hyperscaler practice. A handful of Indian commercial neoclouds are honest enough to use 4 years, which is roughly the longest defensible life given that NVIDIA's Hopper-to-Blackwell-to-Rubin cadence is shipping a new generation roughly every 12–18 months and that H100 spot rental rates have already fallen ~70% from peak.
The reality, as best as I can tell from talking to people who do this professionally and from looking at how lenders price the same hardware (lenders demand 3.5–5 year amortization on GPU-collateralized loans, which is what they actually believe), is that the *real* economic life of an H100 in India is probably **2.5 to 3.5 years**. If you re-do every Indian neocloud's gross margin with that assumption, instead of the 5–6 year assumption that's currently in their books, the numbers compress by 15 to 25 percentage points. A "70% gross margin" Indian AI infrastructure business at 2.5-year economic life is probably a **45–50% gross margin** business. Which is fine — that's a perfectly respectable infrastructure business. It is not, however, a software business. And right now most of these companies are being valued like software.
The IndiaAI providers don't have this problem in the same way, because their P&L is dominated by subsidy receipts rather than chip economics. But when the program tapers — and these programs always taper — the operators who have spent five years optimizing for "win the tender" rather than for "run the chip well" are going to discover, all at once, that the chip they win the next tender against is a Blackwell that does the same workload using a quarter of the power on the same Maharashtra grid. *2027.* Three years from now. Roughly.
---
## 7. What does it actually mean
So what does it mean for you, if you are running an Indian AI startup, or an enterprise AI program at a BFSI firm in Mumbai, or a GCC of a US tech company, or you're just trying to make sense of why every Indian AI deck you see has a 70% gross margin in cell B14?
A few things.
**One:** if you are taking IndiaAI Mission compute, take it. It's free money. But run it on a separate set of books from your commercial workload, and measure your actual cost-per-useful-token as if you were paying ₹249/hr for it, because in 2027 you might be.
**Two:** measure goodput, not uptime. The single highest-leverage FinOps move you can make as an Indian AI operator is to instrument your training and inference workloads against MFU and against $/successful-inference, not against `nvidia-smi`. The cost of doing this in India is comically low — a five-person team is $200K — and the savings, if industry-typical waste figures (~30–50% of AI spend) translate, run into the crores.
**Three:** if you are an investor, the gross margin on the deck is not the gross margin on the business. Re-do it at 3-year depreciation. Re-do it at the commercial GPU rate, not the subsidized one. Re-do it at 30% utilization, which is the realistic frontier number, not 80%, which is the underwriting fiction. The companies that *still* look good after those three adjustments are the real ones.
**Four:** the most undervalued asset in Indian AI right now is **a senior infrastructure engineer who knows what MFU is** and can get a workload from 22% to 40%. That person, in the US, costs $500K. In Bengaluru, that person costs $50K. The arbitrage is enormous, and almost nobody is running it, because the people writing checks have not yet figured out that this is the lever.
The Mission is, on balance, a very good thing. It is buying India a seat at a table the country would otherwise not have a seat at. But it is also, as a side effect, *teaching an entire generation of Indian AI builders that compute is free*, and compute is not free, and the bill is going to come due, and when it does the people who will survive are the ones who have spent the subsidy era pretending they were paying full price.
The cheapest GPUs in the world are not free. They're just being paid for by somebody else, for now.
Welcome to the queue.
---
## Sources
Bedrock AI · Behind the Balance Sheet · Bizety · BW Businessworld · Cast AI · Cerno Capital · CIO Dive · CloudZero · Cybernews · Data Center Dynamics · Deep Quarry · EE Times · Eurostat · Flexera (State of the Cloud 2026) · FinOps Foundation (State of FinOps 2024–2026) · GetDeploying · GMI Cloud · ICICI Direct · IEA Electricity 2026 · Inc42 · Interface · IntuitionLabs · Introl · Jarvis Labs · Latent Space · Levelheaded Investing · Llama 3 technical report · McKinsey · Mercom India · MIT NANDA (GenAI Divide) · neXt Curve · Outlook Business · Oplexa · PaLM technical report (arXiv 2204.02311) · Press Information Bureau · Princeton CITP · RawCompute · SDxCentral · SemiAnalysis (H100 Index, GB200 Benchmarks, Goodput Theory) · Silicon Data · SiliconANGLE · SME Futures · Spheron · Stanford 2025 AI Index · Tech Insider · The Information · theCUBE Research · Thundercompute · Tom's Hardware · Trainy · Varindia · vLLM/PagedAttention paper (arXiv 2309.06180)
---
# Why You Can’t Buy a Data Center Right Now
Source: https://www.timlig.com/posts/ai-supply-chain-crisis/
Published: 2026-02-21
Tags: AI, supply chain, data centers, semiconductors, infrastructure
If you run a data center, or you’re trying to build one, or you’re trying to buy electronic components for one — you already know things are bad. But you might not know *why* they’re bad, or *how* bad, or how long they’ll stay bad. So let me walk you through the whole thing.
The short version is: everyone wants to build AI data centers at the same time, the factories that make the parts can’t keep up, the factories that make the machines that make the parts *also* can’t keep up, and — this is the fun part — you can’t even plug the data centers into the electrical grid because there aren’t enough transformers, and the wait for a new transformer is now **four years**.
Four years! For a transformer! That’s longer than most presidential terms.
-----
## The Basic Problem
Here’s what happened. In November 2022, OpenAI released ChatGPT. It became the fastest-growing consumer app ever. Every big tech company looked at it and said “oh no, we need to build a lot more AI infrastructure, immediately.” And then they all tried to buy the same stuff from the same handful of factories at the same time.
The companies spending the money are the “hyperscalers” — Amazon, Google, Meta, Microsoft, and a few others. In 2023, they spent about $155 billion combined on infrastructure. In 2025, they spent about $443 billion. For 2026, they’ve announced plans to spend somewhere around **$650 to $700 billion**.
That’s a *lot* of money. Amazon alone said it’ll spend $200 billion in 2026. Google doubled to $175 billion. These are numbers that would have seemed insane three years ago, and frankly they still seem insane, but here we are.
The problem is that money doesn’t turn into data centers instantly. You need chips, memory, packaging, servers, power supplies, transformers, cooling systems, optical cables, and construction workers. And right now, almost every single one of those things is in short supply.
-----
## The Seven Bottlenecks
I think of the AI supply chain as a series of choke points. Each one can independently slow everything down, and they all make each other worse. A GPU can’t ship without memory. Memory can’t be assembled without packaging. None of it works without power and cooling. It’s bottlenecks all the way down.
Let’s walk through each one.
### 1. GPUs: The Chip Everyone Wants
NVIDIA makes about 80% of AI chips. Their latest “Blackwell” chips are manufactured by TSMC in Taiwan. They are **sold out through mid-2026**, with a backlog of roughly 3.6 million units. If you want GPU servers, you need to put down a non-refundable deposit 9–12 months in advance.
AMD is the main alternative. They’ve got a credible product now (the MI350), and they scored a huge deal with OpenAI worth potentially $100 billion over four years. AMD’s share of the AI chip market is climbing toward 15–20%. Intel, meanwhile, has basically given up — they cancelled their big AI chip (Falcon Shores) and publicly admitted they’re not meaningfully competing in this market yet.
But even if you can get chips from AMD, you still need everything else on this list.
### 2. HBM Memory: The Real Bottleneck
This one is wild. High Bandwidth Memory (HBM) is the special memory that sits right next to the GPU on the chip. It’s what lets AI chips process huge models quickly. Three companies make it: SK Hynix (62% market share), Micron (21%), and Samsung (17%).
**All three have confirmed their entire production for 2025 and 2026 is already sold.** SK Hynix’s CFO literally said they’ve sold out all of 2026. There is no more to buy.
And the problem gets worse with each new GPU generation, because each one needs *more* memory:
The H100 uses 80 GB. The new Rubin chip will use 288 GB. That’s a 3.6× increase per chip. So even if memory production grows, each chip eats more of it. Memory prices went up **246% in 2025**, and Bloomberg Intelligence says oversupply won’t happen until **2033**.
I’ll say that again: 2033. We’re looking at a structural shortage for the next 7+ years.
### 3. Advanced Packaging: The Quiet Chokepoint
So you’ve got a GPU design, and you’ve got HBM chips. Now you need to stick them together. This is called “advanced packaging,” and it’s done through a process called CoWoS (Chip-on-Wafer-on-Substrate). TSMC basically has a monopoly on this for AI chips.
The situation here is: demand for 2026 is **~1 million wafers**. Available capacity is maybe **120,000–130,000 wafers per month**. TSMC has pre-booked 85%+ of its capacity to its biggest customers. If you’re not NVIDIA, AMD, or a major hyperscaler, good luck getting a slot.
Here’s a detail that captures how tight this is: NVIDIA alone needs about 595,000 CoWoS wafers in 2026 — roughly **60% of total global capacity**. One company. One product.
### 4. Power Transformers: The Show-Stopper
This is the bottleneck that surprises people the most, and honestly it’s the scariest one, because semiconductor supply chains can scale in 1–2 years, but power infrastructure takes **5–10 years**.
A large power transformer — the kind you need to connect a data center to the grid — now has a lead time of **128 to 210 weeks**. That’s 2.5 to 4 years. Before the AI boom, it was 6–8 months.
Why? A few reasons. First, these transformers can’t be mass-produced — each one is custom-designed, individually tested, and weighs 100 to 400 tons. The US has about 10 railcars capable of transporting them. Second, only 20% of large US transformers are made domestically. Third, demand has grown 274% since 2019 while capacity hasn’t remotely kept up. Manufacturers have invested $2 billion in new factories, but most won’t open until 2027–2028.
Wood Mackenzie estimates a **30% supply shortfall** for power transformers in 2025. Prices have gone up 4–6× since 2022.
### 5. Cooling: Air Won’t Cut It Anymore
Old data center racks use 5–10 kilowatts and can be cooled with air conditioning. An NVIDIA GB200 rack uses **120–130 kilowatts**. The next generation (Rubin) will hit 200+ kilowatts. NVIDIA has already shown reference designs at **1 megawatt per rack**.
You can’t air-condition your way out of a megawatt rack. You need liquid cooling — pipes running coolant directly to the chips. Demand for liquid cooling surged 156% year-over-year. Vertiv, a major cooling equipment maker, has a $9.5 billion backlog. CDU (cooling distribution unit) lead times are now 6–9 months.
### 6. Networking: 864 Fibers Per Rack
Every AI rack needs to talk to every other AI rack, very fast. Modern AI clusters use 800G optical transceivers. Each rack requires **864 individual fiber connections**, rising to 1,526 for next-gen systems. McKinsey estimates $150 billion in fiber needs — enough cable to circle Earth 120 times.
China-based manufacturers make about 60% of these transceivers (mainly Innolight and Eoptolink). They’re moving production to Thailand and Vietnam to work around US tariffs.
### 7. Critical Materials: The Ones You’ve Never Heard Of
There are some really obscure bottlenecks that matter a lot. The best example is **T-Glass** — a specialized glass cloth made almost exclusively by one Japanese company (Nittobo). It’s used in the circuit boards inside AI servers. Nikkei Asia called it “one of the biggest bottlenecks for the electronics-making and AI industry for 2026.” Apple, NVIDIA, AMD, and Google have all sent people to Japan to try to secure supply. New capacity isn’t expected until late 2027.
-----
## Where Everything Is Made (And Why That’s Scary)
The AI supply chain is the most geographically concentrated critical infrastructure in the world. Here’s a simplified picture:
I want to stress how unusual this is. **Ninety percent** of the world’s most advanced chips are made on one island. One Dutch company makes 100% of the machines required to produce them. One Korean company makes 62% of the memory. If any single link in this chain breaks — whether from a natural disaster, a geopolitical crisis, or just a factory fire — the entire global AI buildout stops.
This is why there’s so much government money flowing into “chip sovereignty.” The US CHIPS Act has put $30.9 billion into domestic semiconductor manufacturing. TSMC is building a $165 billion campus in Arizona — the largest foreign investment in US history. The EU has its own Chips Act that’s attracted €69 billion. But new fabs take 3–5 years to build and cost $15–20 billion each, so relief is years away.
-----
## The Geopolitics: Trade War Meets Chip War
The US and China have been in an escalating chip war since October 2022, when the Biden administration first blocked exports of advanced AI chips to China. Here’s the short timeline:
- **Oct 2022**: US blocks A100/H100 exports to China
- **Oct 2023**: Rules expanded to close the H800 loophole
- **Jan 2025**: Biden’s “AI Diffusion Rule” creates a 3-tier global licensing framework
- **Apr 2025**: Trump bans even the “compliant” H20 chip; NVIDIA takes a $5.5 billion write-off
- **Jul 2025**: Trump rescinds the Diffusion Rule, creates a new AI Action Plan
- **Aug 2025**: H20 and AMD MI308 approved for China with a novel 15% revenue-sharing deal
Meanwhile, China has hit back with export controls on critical minerals — germanium, gallium, and rare earths that are essential for chip manufacturing. They’ve also poured $138 billion into their own state-backed semiconductor fund. But they’re still far behind: China’s AI chip output is estimated at **only 1–4% of US production**.
Sovereign AI is also creating new demand centers. The UAE and Saudi Arabia are planning $100+ billion in AI infrastructure. The US Commerce Department recently authorized 70,000 NVIDIA GB300 chips for the Middle East. Basically, everyone wants AI chips, and there aren’t enough for everyone.
-----
## The Power Problem Is the Worst One
I’ve saved the most important section for last, because I think this is the part most people underestimate.
You can speed up chip manufacturing. TSMC is expanding CoWoS capacity. Samsung is boosting HBM production 50% for 2026. The semiconductor industry is really, really good at scaling up when there’s demand.
But power infrastructure? That runs on a different clock entirely.
Goldman Sachs projects data center power demand will grow **165% by 2030**. The IEA says data center electricity consumption will double to ~945 TWh by 2030 — about 3% of all electricity on Earth.
Northern Virginia, the world’s biggest data center market, already consumes over 25% of the state’s electricity. Dublin crossed 10% and imposed a moratorium. Amsterdam capped new data center growth until 2030. Over 100 US counties and cities have passed temporary freezes on new data centers.
The hyperscalers are desperate enough to go nuclear. Microsoft signed a 20-year deal to restart **Three Mile Island** (yes, that Three Mile Island). Amazon secured 1.92 GW of nuclear capacity. Meta and Google have their own nuclear deals. But nuclear restarts take years, and small modular reactors won’t deliver meaningful power until the 2030s.
Here’s the bottom line for procurement people: **you now need to plan power infrastructure 2–4 years in advance**. If you haven’t ordered your transformers yet, you’re already behind.
-----
## So When Does It Get Better?
Honestly? Not soon. Here’s my rough outlook:
**CoWoS packaging** is the most likely to ease first. TSMC is aggressively expanding capacity and expects to roughly triple output by end-2026. The acute shortage will ease, though it’ll remain tight.
**GPU availability** should improve in late 2026 as AMD scales up and TSMC brings more capacity online. But demand is also accelerating, so “improve” means “slightly less impossible,” not “easy.”
**HBM memory** stays structurally tight through at least 2028. All three manufacturers are expanding, but demand is growing just as fast.
**Power infrastructure** is the long pole. Transformer manufacturing capacity won’t meaningfully expand until 2027–2028. Grid connections in major markets are measured in years. This is the constraint that will define the AI era.
-----
## What Does This Mean for You?
If you’re in data center procurement, here’s the actionable version:
**Plan 2–4 years ahead for power.** If you need a transformer or a new grid connection, you should have ordered it yesterday. The lead times are not going to shorten meaningfully before 2028.
**Plan 12–18 months ahead for GPU/server allocations.** Put down deposits early. The companies that secure supply first will be the ones that can actually build.
**Expect elevated pricing through at least 2027.** Memory up 246%. Transformers up 4–6×. Cooling equipment backlogs growing. This is the new normal, not a temporary blip.
**Diversify suppliers where possible.** AMD’s MI350/MI450 are real alternatives to NVIDIA now. Ethernet networking is overtaking InfiniBand. Multiple cooling vendors are scaling. Don’t be single-sourced on anything.
**Watch the geopolitics.** US-China chip export rules change every few months. Tariffs on Chinese-made optical transceivers may force supply chain reshuffling. CHIPS Act subsidies are available but time-limited.
-----
## The Big Picture
Previous chip shortages — the 2020-2021 COVID crunch, the 2011 Thailand floods — resolved within 12–18 months as demand normalized. This one is different, for three reasons:
1. **Demand is exponential.** AI compute is growing ~2.25× per year. It’s not going to normalize — it’s accelerating.
1. **Manufacturing is maximally concentrated.** One foundry, one lithography company, three memory makers. There’s no slack in the system.
1. **Infrastructure runs on decade timescales.** You can scale chip production in 2 years. You can’t scale the power grid in 2 years.
The combination of exponential demand growth, extreme geographic concentration, and decade-long infrastructure timescales is what makes this a structural crisis rather than a cyclical one. The industry is spending more money than ever, building faster than ever, and it’s still not enough. It won’t be enough for a while.
The companies that figured this out early — that locked in transformer orders in 2023, secured HBM allocations in 2024, and started building relationships with cooling equipment makers before everyone else realized they’d need liquid cooling — those are the companies that will have working data centers in 2026 and 2027. Everyone else is in the queue.
Welcome to the queue.
-----
*Data sourced from Bloomberg Intelligence, Goldman Sachs Research, McKinsey, CNBC, Fusion Worldwide, POWER Magazine, Dell’Oro Group, Deloitte, Congressional Research Service, TrendForce, NVIDIA, SK Hynix, and other industry sources cited throughout. Data current as of February 2026.*
---
# Why Timlig Exists For A Minute
Source: https://www.timlig.com/posts/why-timlig/
Published: 2025-12-06
Tags: meta, naming, vibe
Timlig is just a stray keyboard mash that happened to be available. No naming workshop, no brand deck. It is the digital equivalent of grabbing whatever sticky note is on the desk and scribbling a thought before the train door closes.
The point is to keep reminding myself how much of our work is temporary anyway:
- We spin up feature branches, throw them away, and only merge the bits that survive daylight.
- We build prototypes that exist just long enough to prove a point, then delete the repo without ceremony.
- We rename services every quarter because the domain shifted and nobody wants to maintain the old story.
If you have ever dumped a log with `mktemp` or copy-pasted a hacky script into `/tmp`, you already get the vibe. Timlig is that: a placeholder handle for shipping thoughts before they evaporate.
So yeah, the name is random on purpose. The site might move, the CSS might change, and half these ideas might get rewritten. That is fine. The goal is to keep publishing while the water is still warm instead of waiting for a forever home.