man talking on the phone

The Missed Call Era of Agentic AI and What it Means for Healthcare Leaders

|

|

11–16 minutes

read

We are in the missed call era of hospital AI — a period of genuine but temporary scarcity, where the infrastructure is real and expensive and constrained, and where smart teams are rationing it with creative workarounds. This will pass. The hospital leaders who learn to work well under today’s constraints will extract disproportionate value…

A few weeks ago, mid-workflow, Claude stopped. I had hit my usage limit.

My first instinct was irritation. My second was recognition.

I had felt this exact thing before — not with AI, but with a Nokia 3310, a prepaid SIM, and a friend I needed to reach on the other side of the city. Though the message from Claude was different in form, it was identical in logic.

I use AI tools heavily — Claude, Gemini, ChatGPT — across research, writing, analysis, and client work. Usage limits, for me, are a recurring operational reality. I have learned to work around them with some ingenuity.

But this particular instance made me stop, and made me think about what it means for the hospitals I work with. Because if I — a single user doing knowledge work — am already rationing AI capacity, the hospital that has deployed an AI-powered post-discharge follow-up system contacting hundreds of patients per month has an orders-of-magnitude larger version of the same problem. And the hospital that has not yet deployed such a system, and is holding back partly because of cost and reliability concerns, is living with a different version of the same constraint.

The Missed Call

If you were using a mobile phone in India between 2000 and 2010, you know what a missed call was. Not an accidental one. A deliberate ring, then an immediate hang-up before the other party could answer. Zero rupees spent. Full message transmitted.

I’ve reached the hospital. Come to the gate. Call me when you can — I’m out of balance.

The missed call was ingenious because it had to be. Talk time was expensive. Both outgoing and incoming calls cost money. The infrastructure — towers, spectrum, interconnects — had consumed enormous capital to build, and that capital was being recovered through per-minute billing. Scarcity was not a policy choice. It was an economic reality.

And so people adapted. A parallel communication layer grew up alongside the formal one, running on zero cost. Missed calls acquired meanings. Patterns of missed calls acquired more meanings. It was improvisation at civilisational scale.

We are doing it again — but this time with agentic AI.

If you are deploying agentic AI tools in your hospital, your teams have probably already started the equivalent. Someone has figured out that batching lower-acuity patient follow-ups early in the month preserves the quota for high-risk discharges that arrive later. Someone has built a prompt library to keep AI-generated patient communications as short as possible, conserving tokens. One department has quietly designated specific functions — discharge summaries — for AI, and continued doing patient calls by phone, preserving the AI allocation for documentation where it saves the most time.

These rationing responses to agentic AI usage limits are like missed calls. Ingenious responses to a scarce and expensive resource.

The scarcity is real. Running large language models — the AI systems that power these tools — requires significant compute infrastructure. Inference is not free. API quotas are not an arbitrary inconvenience. They are the per-minute billing of 2026.

But the missed call was a phase, not a permanent condition.

What agentic AI looks like in an Indian hospital right now

Before examining the infrastructure arc, it is worth being precise about what kind of AI we’re talking about.

Agentic AI refers to AI systems that take actions on behalf of an organisation — not just answer questions, but do things. The distinction matters. A doctor using ChatGPT to draft a discharge summary is using a generative AI tool. An AI system that detects which patients are due for discharge in the next 24 hours, generates draft summaries for each, routes them for doctor sign-off, and sends structured clinical communications to referring GPs upon discharge — without being prompted each time — is an agentic system.

Agentic AI is slowly being deployed in Indian hospitals, though the term is rarely used. What hospital leaders are calling “AI-powered follow-up,” “ambient scribing,” and “automated TPA follow-up” are, structurally, agentic systems.

Four categories are in popular use, and more will likely follow.

  1. Patient communication. The most deployed category in India. Appointment reminders, post-discharge follow-up, discharge instruction delivery, and chronic disease check-in calls are running at hospitals in Tier 1 cities and, increasingly, in Tier 2. In March 2025, Apollo Hospitals — which operates over 10,000 beds across India — announced it was increasing AI investment to automate routine clinical and administrative tasks, one of the clearest signals from a large hospital network that agentic systems have moved from experiment to operational priority. The constraint limiting full deployment is not willingness — it is the quota problem I described above. High-volume follow-up programmes hit usage limits mid-month. The workaround, for now, is batching and triage.
  2. Clinical documentation. Ambient scribing — AI that listens to OPD consultations and generates a structured clinical note — has moved from pilot to limited rollout at several urban private hospitals. AI-assisted discharge summary generation is further along. The constraint here is medico-legal: clinical documentation carries accountability, and hospital leadership is appropriately cautious about the sign-off workflow. Most deployments require a physician to review and approve every AI-generated document, which slows throughput but reduces risk.
  3. Administrative workflow. Prior authorisation and TPA/insurance follow-up are natural targets for automation — high-volume, rule-based, and currently consuming significant staff time. Several platforms now offer AI-assisted TPA follow-up built specifically for the Indian insurance environment, including Ayushman Bharat claims processing. The constraint is integration: these systems need to connect with the hospital’s existing HIS, and HIS integration in Indian hospitals remains fragmented.
  4. Clinical decision support. The most nascent category. Triage flagging, diagnostic assistance in radiology, and risk stratification in ICUs are real and active in research and premium care settings. The constraint here is not just economics — it is the genuine complexity of clinical accountability. Who is responsible when an AI-flagged diagnosis is wrong? The regulatory and medico-legal framework in India has not yet settled this question.

Why the constraints are temporary

The arc that produced the missed call’s obsolescence is not unique to mobile telephony. It is the standard trajectory of infrastructure technology.

Large investment in new infrastructure → genuine scarcity and high unit cost → rationing and workarounds → competition enters → scale drives unit economics down → commoditisation → abundance.

This arc applies to electricity, internet connectivity, cloud computing, and mobile telephony. In each case, the infrastructure required significant upfront capital, the early cost to the end user reflected that capital recovery, and sustained competition eventually made the service cheap enough to treat as a background assumption of doing business. The endpoint was not just affordability — it was invisibility. The businesses that understood the arc early and built their operations assuming the endpoint, not the current midpoint, consistently outperformed those that waited.

The same arc is in motion for AI compute.

The infrastructure investment is documented. OpenAIGoogle DeepMindAnthropicMicrosoft, and Meta have each committed tens of billions of dollars to AI infrastructure over the past four years. The capital expenditure required to train and serve large language models is the reason API costs are as high as they are — and the reason they have been falling as fast as they have.

Since GPT-4 launched in early 2023, the cost of running a model with equivalent performance has fallen by more than 97%. A million tokens processed by GPT-4 at launch cost approximately $30. The same output from a current model of equivalent capability costs under $1. That is a collapse, not a reduction — driven by three forces: specialised AI inference chips that process tokens far more efficiently than general-purpose GPUs; model efficiency improvements, including smaller and better-optimised models that match the performance of earlier large models at a fraction of the compute; and open-source pressure from models like Llama and Mistral, which have forced commercial providers to compete aggressively on price.

The India-specific version of this story is not just an analogy. It is a lived reference. Reliance Jio‘s entry in 2016 did not invent the infrastructure arc — the arc was already in motion. But it collapsed the timeline. Data that cost ₹250 per GB in 2016 cost ₹10 per GB three years later and effectively nothing thereafter. Hospital leaders who built patient communication systems on the assumption that mobile internet would remain expensive made a costly error in strategic judgment. Those who assumed the arc would complete — and built for the endpoint — benefited disproportionately.

The AI compute arc will complete. The question is not whether hospitals should build on AI infrastructure. It is whether they should start now — under constraints, learning as they go — or wait until the economics are perfect. The hospitals that waited for Jio to happen before investing in mobile-based patient communication were two years late to an advantage that compounds.

What the missed call era teaches a hospital leader

The missed call is not only a metaphor. It contains two practical lessons for hospital leaders operating under the current constraints.

Principle 1: The constraint is a teacher.

The people who built the most efficient communication habits during the high-tariff era were not just making do. They were developing a discipline. The constraint forced them to think about what information actually needed to move, how to compress it, what could wait and what couldn’t. When unlimited plans arrived, those habits did not become useless — they became multiplied. The communicators who had learned precision were better communicators when bandwidth became free. The same logic applies to hospital AI.

The teams building AI-assisted workflows under today’s constraints are learning something that cannot be bought when the constraints lift: what is actually worth automating.

A post-discharge follow-up system running under a token budget teaches the team which patients most need a follow-up call, which cases can be handled with a shorter automated message, and which require a human. An AI-assisted documentation programme run under a quota teaches the clinical team which documents benefit most from AI generation and which are so physician-specific that automation adds friction rather than reducing it. A TPA follow-up tool with monthly usage limits forces the billing team to triage claims by value and urgency — a discipline that has value independent of AI.

This institutional knowledge — what to delegate, what to supervise, what context the AI needs to be useful, where AI fails and why — is the durable asset. When costs fall and quotas lift, the hospital that has run constrained workflows for eighteen months will scale from experience. The hospital that waited for perfect economics will begin where the other hospital was in month one.

The constraint, used correctly, is not a problem. It is learning.

Principle 2: Do not design your AI strategy for permanent scarcity.

A hospital in 2012 that concluded mobile phones were too unreliable to build patient communication systems around was not being conservative. It was making a prediction about technology’s direction of travel — and it was wrong. The infrastructure cost was being recovered. Competition was accelerating the recovery. The hospital’s caution was, in effect, a bet against a well-documented historical pattern.

The equivalent error today is treating today’s API quotas as permanent and making strategic decisions accordingly: keeping AI adoption minimal because cost-per-interaction is too high at current volumes; avoiding workflow integration because reliability is imperfect; concluding that agentic systems are a “future investment” rather than a current priority.

These conclusions may be tactically reasonable for specific use cases at specific price points. But as a general orientation — AI dependency is risky because today’s constraints are real — they reflect the wrong model of how infrastructure technology works.

The specific version matters. A hospital that examines a post-discharge AI follow-up system, finds the monthly cost at current quotas to be ₹40,000, determines that this exceeds the cost of the staff time it replaces, and therefore passes — has done a sound unit economics calculation for today’s costs. It should also note, in that same analysis, the cost trajectory. If that ₹40,000 becomes ₹15,000 in eighteen months — a realistic projection given the documented pace of AI inference cost reduction — the decision changes. The hospital that has already built and tested the workflow at ₹40,000 will be running it profitably at ₹15,000. The hospital that waited will be starting from scratch.

A practical framework: AI adoption tiers

The question of which AI investments to make now, which to pilot, and which to defer is a unit economics question — but one that needs to be made against a three-year projection, not today’s cost sheet.

Three tiers organise the decision.

Tier 1 — In Use Today

Use cases with clear positive ROI at current costs. High-volume, low-stakes, well-understood. The key test: even at today’s API pricing, is AI cheaper than the staff time it replaces?

  • Automated appointment reminders and confirmation come to mind. A hospital doing 300 OPD appointments per day spends significant staff time on reminder calls, confirmation tracking, and no-show management. AI handles this at a fraction of that cost, consistently, regardless of front-desk workload.
  • AI-assisted discharge summary generation with physician sign-off is another good one; the time savings on documentation are immediate and measurable.
  • TPA and insurance follow-up automation may qualify for hospitals with significant insurance volumes, where the repetitive nature of follow-up makes it exactly the kind of work AI handles well.

Tier 2 — Pilot with attention to unit economics

Use cases where the value is clear but the cost-benefit at current pricing requires monitoring. Deploy at sufficient scale to build institutional knowledge; do not commit to full rollout until the pricing improves.

  • Ambient scribing in OPD consultations sits here. The clinical value — more accurate notes, less documentation burden on physicians — is real. The cost per consultation at current AI pricing, combined with the workflow integration required, makes full-scale rollout economics marginal at most Indian hospitals. Pilot it in one department, measure actual physician time savings, track the cost per consultation, and set a pricing threshold at which you scale.
  • AI-assisted triage protocols and Level 1 patient inquiry handling — basic clinical questions routed to AI before reaching clinical staff — are also Tier 2: promising, worth learning, not ready for full commitment.

Tier 3 — Wait and watch

Diagnostic AI with direct clinical accountability, autonomous patient conversations beyond structured queries, and AI-led clinical decision support in high-stakes specialties. The clinical value may be real; the medico-legal framework in India is not yet ready, and the reliability requirements exceed what current systems deliver consistently.

The quiet end

The missed call died without ceremony.

Nobody announced it. There was no moment when Indians collectively decided the ritual was over. One day, Jio arrived. Plans got cheaper. Then free. And the missed call — that intricate, compressed, civilisational improvisation — simply stopped being necessary.

I feel that is what is coming for agentic AI quotas and token limits. A time will come when we’ll stop making these “missed calls”.

The hospitals building AI workflows today — rationing tokens, batching requests, building prompt libraries, learning which use cases are worth automating and which are not — are not just making do. They are accumulating the institutional knowledge that will determine who extracts the most value from AI when those limits no longer exist.

Until then: be ingenious.

How AI and technology adoption fits within the broader framework of hospital operations — from process design to workflow automation — is covered in the operational excellence guide for private clinics and hospitals in India.


If this was useful, there’s more where it came from.

I’m Aviral. I help Indian healthcare organisations grow and run better, by putting the right systems in place. Subscribe to stay updated.

Leave a Reply

Discover more from Aviral Prakash

Subscribe now to keep reading and get access to the full archive.

Continue reading

I write about the business of medicine - how healthcare practices get built and run better.

Subscribe if this is useful to you.