Ads_970x250

Medical AI Matched the Doctors. Are India's Hospitals Still Saying No?

Two studies show AI systems handling simulated cases at the physician level. Their researchers and Indian healthcare leaders still want a clinician responsible for every consequential decision.

Topics

  • Image Credit- Chetan Jha/ MIT Sloan Management Review India

    Key Takeaways

    01

    MIRA and AMIE performed at or above physician level on several measures, but neither system was tested prospectively with patients in clinical care.

    02

    A strong benchmark cannot show how a system will cope with incomplete records, unpredictable conversations, local workflows, or the consequences of a wrong decision.

    03

    India has begun building national testing and governance infrastructure for health AI. Hospitals still need their own rules for authorization, monitoring, and accountability.

     At Yashoda Hospitals in Hyderabad, Dr. Sachin Marda often corrects a misconception about robotic surgery. Some patients think the robot operates. It does not. Marda controls every movement.

    He sees medical AI in much the same way. A system can assemble evidence, flag a risk or suggest what a clinician might investigate next. Marda is not ready to let it order a test or choose treatment without a doctor. “Medicine is far more than pattern recognition and prediction,” he said.

    Two studies published in Nature on June 17 showed how much more capable these systems have become. One followed an AI agent through a simulated emergency department visit, from taking a history to prescribing drugs and planning an admission. The other tested whether an AI system could manage a condition over several consultations and adjust its recommendations as the case changed.

    The results compared well with those of doctors. MIRA, or Medical Intelligence for Reasoning and Action, reached 87.8% diagnostic accuracy in a head-to-head test. Four board-certified physicians scored 78.1%, while a separate group of six physicians with mixed levels of experience scored 71.1%. Google’s AMIE, the Articulate Medical Intelligence Explorer, was rated non-inferior to 21 primary care physicians on management reasoning and produced more precise investigation and treatment plans.

    Neither system treated a patient. MIRA worked in a sandboxed electronic health record, while AMIE spoke with trained patient actors in a virtual examination. Both research teams said more evidence would be needed before clinical use. Indian hospital leaders now have to decide how much freedom to give these systems. They can resemble a doctor in a study, but none has ever carried responsibility for a patient.

    The Studies Stop Short of the Bedside

    MIRA was tested on more than 500 cases drawn from MIMIC-IV, a database of deidentified records from a Boston hospital. It could use 11 electronic health record tools and choose among more than 85,000 possible actions. It interviewed a simulated patient, ordered and interpreted tests, generated diagnoses, prescribed medications, scheduled procedures and recommended admission.

    The comparison with physicians covered 311 matched cases across eight conditions. MIRA performed particularly well on pancreatitis, reaching 95.2% diagnostic accuracy, compared with 78.6% for the board-certified physicians. The model and both physician groups found pneumonia and urinary tract infections more difficult.

    MIRA requested more of the blood analytes found in the reference records than the doctors did. It did not order more expensive imaging. Its overall test selection also remained below the level recorded in the original hospital data. A hospital considering such a system would still need to watch resource use separately from diagnostic accuracy. The two do not move together automatically.

    The prescribing results were similarly encouraging, with limits that are easy to miss in the headline figure. Of 468 prescriptions across 56 cases, 467 contained free-text dosing instructions that a reviewer judged relevant, useful and correct. Route of administration, the weakest field, was correct in 97% of prescriptions.

    These were retrospective cases presented in a controlled interface. The simulated patient answered from information already recorded in the case history, producing conversations that the authors acknowledged were probably more orderly than those in an emergency department. MIRA and the doctors encountered the same simulated patient, which made the comparison fair. It did not make the conversation realistic.

    The researchers also could not rule out some overlap between the widely used MIMIC-IV data and material seen during model training. They wrote that the reported performance might represent an upper bound and could overstate how well MIRA would generalize to unfamiliar cases.

    Their recommendation was cautious. Early clinical use, they suggested, should concentrate on bounded tasks such as reconciling medications, preparing laboratory panels and drafting consultation requests, with a physician reviewing the work. Medical AI agents, the authors wrote, should not currently be designed to replace healthcare professionals.

    The AMIE study reached a similar conclusion from a different type of test. AMIE managed 100 cases over three visits in five specialties. The system and the physicians could consult a collection of clinical guidelines as a patient’s symptoms, test results and response to treatment changed.

    The physicians and patient actors were based in India and Canada, and clinical providers in the two countries prepared the scenarios. The guideline reference standard, though, came from the UK’s National Institute for Health and Care Excellence and BMJ Best Practice. The authors said the case mix did not resemble a normal clinical practice and that patient actors could not reproduce actual care. Indian participation made the experiment broader. It did not amount to validation in an Indian hospital.

    AMIE’s own researchers described it as an “art-of-the-possible” demonstration. They said the system was not ready for clinical care and would need prospective studies involving patients, along with appropriate ethical and safety oversight.

    “Medicine is far more than pattern recognition and prediction.” 

    Dr. Sachin Marda, Clinical Director and Senior Consultant Oncologist and Robotic Surgeon, Yashoda Hospitals, Hyderabad

    A Test Order Sets Several Systems in Motion

    Cancer care explains Marda’s reluctance to hand over clinical authority. A recommendation for a PET scan, a biopsy or a genetic test can alter the treatment plan, add considerably to the bill and alarm a family. Patients with the same cancer stage may need different care because of their age, other illnesses, physical condition, molecular profile, family support or wishes.

    Recognizing a disease is only one part of that decision. The physician must decide whether another test is useful, what its result would change and whether the patient can bear the procedure and its cost.

    The same operational complications appear in home healthcare. Bharadwaj J.S.R., head of IT and products at Apollo Home Healthcare, said a test order may require consent, a home visit, a caregiver, sample transport, laboratory coordination and follow-up by the treating doctor. It also carries financial and legal consequences.

    An AI system may use care pathways and a patient’s history to recommend an investigation, Bharadwaj said. The approval should come from a physician or another authorized clinician. “AI should become the smartest clinical assistant in the room before it becomes the decision maker,” he wrote.

    “AI should become the smartest clinical assistant in the room before it becomes the decision maker.”

    —  Bharadwaj J. S. R., Head of IT and Products, Apollo Home Healthcare

    Apollo Home Healthcare examines whether a tool fits existing clinical processes, works consistently across patient groups and saves clinicians time. The company also considers whether doctors can understand the basis of a recommendation. A high score has limited value if the tool adds work, sits apart from the clinical system or loses the confidence of its users.

    “In home healthcare, reliability matters more than novelty,” Bharadwaj said.

    Before hospitals allow a system to order tests independently, he added, they would need evidence from Indian populations, a clear account of what the system is authorized to do and an answer to who bears responsibility when its recommendation causes harm.

    A Model’s Score Slips Once Real People Are Involved

    The MIRA and AMIE experiments were designed to make reliable comparisons possible. Real conversations introduce problems that a benchmark cannot neatly control. Patients forget medicines, confuse dates, omit earlier diagnoses, or describe the same symptom differently depending on language and stress.

    A randomized study in Nature Medicine tested what happened when members of the public used language models for medical advice. The study involved 1,298 adults in the UK and ten written medical scenarios.

    Given the complete scenarios directly, the models identified relevant conditions in 94.9% of cases. Participants using the same models identified them in fewer than 34.5% of cases. They chose the correct level of care in fewer than 44.2% of cases, no better than participants using their usual sources, which were mostly internet searches or their own knowledge.

    The models had access to the relevant medical knowledge. The conversations often failed to draw it out. Some participants left out important information. Models sometimes misunderstood questions. Users also received mixtures of good and poor advice and did not reliably act on the useful parts.

    The study concerned members of the public, not clinicians working inside a hospital. Its relevance lies in the difference between evaluating a model alone and evaluating the people and system around it. A hospital cannot infer the quality of an AI-assisted consultation from the model’s standalone score.

    Alok Katiyar, co-founder of WeClinic Homeopathy, applies a practical test to new systems: Does the tool fit the consultation, protect patient information, save the physician time and improve the patient’s experience?

    “Novelty alone has never been a good enough reason to adopt technology in healthcare,” he said.

    Katiyar sees the most immediate value in preparing for the consultation. AI can organize records, trace symptoms across visits and place the relevant history in front of the doctor. “The first job of AI is to give doctors their time back,” he said.

    Errors Can Arrive Through the Medical Record

    AI safety discussions often concentrate on information a model invents. A model can also repeat false information that it has been given, particularly when the information looks authoritative.

    Researchers tested that weakness in a study published in The Lancet Digital Health. They ran more than 3.4 million prompts across 20 language models using fabricated medical claims drawn from social media discussions, simulated clinical cases and real hospital discharge notes altered to include one false recommendation.

    The researchers first tested 158,000 prompts without any added rhetorical framing. Models accepted the fabricated material in 31.7% of them. The rate rose to 46.1% when the false recommendation appeared inside a hospital note. For misinformation presented in the style of a social media discussion, it was 8.9%.

    Clinical language appeared to lend the false claim credibility. A system linked to an electronic health record could therefore retrieve an old mistake, a faulty instruction or an incorrect diagnosis and present it as part of a coherent recommendation.

    The study did not test a complete hospital deployment with retrieval controls, local guidelines and clinician review. It did identify a weakness that hospitals can test for. Checking the quality of the source record cannot end when implementation begins. New notes, corrections and inconsistencies will continue to enter the system.

    Dr. (Brig.) Ranjit Ghuliani, medical superintendent at Noida International Institute of Medical Sciences College and Hospital in Greater Noida, said the institution had reviewed AI products that performed well in research but were not introduced into routine care. Concerns included limited validation in Indian settings, poor integration with hospital systems, security risks, uncertain responsibility and doubts about whether performance would hold up after deployment.

    AI can assist with documentation, data analysis and pattern detection, Ghuliani said. “Diagnosis and treatment should remain with practicing doctors.”

    Fragmented records add another difficulty. A patient’s history may be divided among hospitals, laboratories, pharmacies and home-care providers. One institution may record information in structured fields while another stores it in free text or on paper. The model can only reason from the information it receives.

    “The first job of AI is to give doctors their time back.”

    — Alok Katiyar, Co-founder, WeClinic Homeopathy

    THE EVIDENCE BASE

    Two June 2026 studies in Nature mark the strongest showing yet for autonomous medical AI. MIRA reached 87.8% diagnostic accuracy in a head-to-head against physicians, and Google’s AMIE was rated non-inferior to 21 primary care physicians at managing cases over several visits. Both teams warned against using the systems to replace clinicians.

    Two further 2026 studies map the risks. An Oxford-led trial of 1,298 people found that users working through the models identified conditions far less accurately than the models did alone. A Mount Sinai study in The Lancet Digital Health found the models accepted a fabricated medical claim 46.1% of the time when it was embedded in a realistic hospital note.

    India Has Begun Building the Rules

    India established two pieces of national health AI infrastructure in February. The Ministry of Health and Family Welfare launched the Strategy for Artificial Intelligence in Healthcare for India, or SAHI, and the Benchmarking Open Data Platform for Health AI, or BODH.

    SAHI guides evaluating, adopting and integrating AI across public and private healthcare. It covers governance, ethics, data quality, interoperability, workforce readiness and system security.

    BODH, developed by IIT Kanpur with the National Health Authority, is intended to test models on diverse, anonymized health data. Its assessments cover performance, robustness, bias and the ability to generalize before deployment at population scale.

    In May, IndiaAI and the Indian Council of Medical Research signed an agreement linking computing infrastructure with biomedical research. ICMR is to contribute anonymized, ethics-approved datasets, models and toolkits from its Medical Information Data for AI Solutions framework to the AIKosh platform. IndiaAI will provide subsidized access to computing resources. The organizations also plan to develop applications around Indian public-health priorities.

    Existing rules cover part of the field. The Central Drugs Standard Control Organization says AI-enabled medical devices fall within the Medical Devices Rules, 2017. Applicants for high-risk devices must provide technical documentation, software verification and validation, risk controls, clinical evidence and quality-management records. The ICMR’s 2023 ethical guidelines address consent, governance and the responsibilities of developers, clinicians, institutions, sponsors and ethics committees.

    None of these measures gives a general-purpose AI system the authority of a licensed clinician. Nor do they settle how responsibility would be divided among a doctor, hospital, software vendor and model provider after a harmful recommendation. Hospitals need to resolve those questions before a pilot reaches routine care.

    What Leaders Should Settle Before Deployment

    Hospital Executives and Boards

    Begin with the action the system will be allowed to take. Summarizing records carries a different risk from recommending a diagnosis, and recommending a diagnosis differs again from placing an order.

    Require a named clinician to approve diagnoses, prescriptions, investigations and treatment decisions. Ask who is responsible for reviewing the recommendation, how quickly the review must happen and what the system should do when no reviewer is available.

    Investment decisions should include integration costs and the effect on clinical workload. A model that creates another inbox, another login or another set of alerts may consume more time than it saves.

    Clinical and Quality Leaders

    Test the combined performance of the system and its users. Give clinicians incomplete histories, contradictory records, unusual presentations and patients who describe symptoms poorly. Measure whether the system changes their decisions, not only whether they like its output.

    Review false positives, missed conditions, unnecessary tests and failures to escalate care. Track performance by patient group and clinical setting. A hospital-wide average can hide a serious weakness in one specialty or population.

    Create a clear route for clinicians to reject, correct and report an AI recommendation. Their corrections should feed into regular safety reviews.

    Technology and Data Leaders

    Record the source of every recommendation and every action taken after it. Access controls, audit trails and version records should make it possible to reconstruct what the system saw, what it produced and who approved the result.

    Stress-test the system with errors already present in medical records. Monitor whether it repeats outdated diagnoses, conflicting drug instructions or copied text from earlier visits.

    Model updates deserve fresh validation. Performance established for one version, dataset or workflow should not be assumed to carry over to the next.

    Medical AI has become capable enough to perform several linked clinical tasks in a controlled setting. That is a real technical achievement, but it does not amount to a transfer of clinical authority.

    For now, the practical division of labor is clear. The system can prepare the case, retrieve evidence, compare options, and suggest a course of action. A clinician remains responsible for what happens next.

    RESEARCH CONTEXT

    This article draws on written responses from Bharadwaj J.S.R., Head of IT and Products, Apollo Home Healthcare; Alok Katiyar, Co-founder, WeClinic Homeopathy; Dr. (Brig.) Ranjit Ghuliani, Medical Superintendent, Noida International Institute of Medical Sciences College and Hospital, Greater Noida; and Dr. Sachin Marda, Clinical Director and Senior Consultant Oncologist and Robotic Surgeon, Yashoda Hospitals, Hyderabad. Principal sources include the 2026 Nature studies on MIRA and AMIE, the Oxford-led Nature Medicine study on AI chatbots for medical advice, the Lancet Digital Health study on medical misinformation, a 2026 review indexed by the US National Library of Medicine on whether AI will replace physicians, and recent announcements from India’s Ministry of Health and Family Welfare, including SAHI, BODH, the IndiaAI-ICMR agreement, ICMR’s ethical guidelines for medical AI, and reporting from the India AI Impact Summit 2026.

    Read next: The Transformation Paradox — Why Organizational Readiness, Not Technology, Determines Whether Strategy Survives Disruption

    Topics

    More Like This

    You must to post a comment.

    First time here? : Comment on articles and get access to many more articles.