How AI intake changes incident reporting
Short answer
AI intake swaps a long form for a conversation. The reporter describes what happened, and the system asks follow-ups and drafts the fields for a person to review. The risks are over-labeling and invented causes. Studies test labeling accuracy, not report volume, so a person must still sign.
AI intake swaps blank fields for a conversation
A form asks the reporter to turn an event into fields. AI intake lets the reporter describe it in their own words. It asks only the follow-up questions still unanswered, and fills the fields as a draft for a person to review.
| Dimension | Form | AI conversation |
|---|---|---|
| Starting point | Blank fields in the designer's order | "What happened?" in the reporter's own words |
| Who structures the data | The reporter | The system drafts; a person confirms |
| Missing details | Often blank or "N/A" | Follow-up questions ask for them |
| Typical failure | Skipped fields and one-word answers | Over-labeling, invented causes, leading questions |
Forms are predictable, need no model, and suit fixed fields such as OSHA Form 301. AI intake earns its place where the story is the useful part of the report, and the reporter dislikes forms.
The evidence is thinner than vendors imply
We found studies of AI models labeling and summarizing incident reports, and one small voice pilot. We found no peer-reviewed study showing that conversational intake raises reporting rates or report quality. Each study used a specific model, so read the numbers as signals.
| Study | What was tested | What it found |
|---|---|---|
| Johnson and colleagues, Joint Commission Journal on Quality and Patient Safety, 2024 | GPT-3.5 labeling 370 obstetric incident reports, checked by a clinician | Sensitivity 85.7%, specificity 97.9%. The model applied 79 labels to the reviewer's 49 (positive predictive value 53.2%). The reviewer approved its reasoning for 60.8% of labels. |
| Denecke and Paula, Studies in Health Technology and Informatics, 2025 | Gemma-2 extracting events, causes and contributing factors from 10,063 Swiss incident messages; 100 checked by hand | Events 92% accurate, causes 84%, contributing factors 72%. Errors came from hallucinating or over-interpreting. |
| Dobranski and colleagues, International Journal of Radiation Oncology, Biology, Physics, 2026 | Local models summarizing and tagging radiation oncology incident reports at two cancer centers; 600 expert ratings | Ratings improved between rounds. In round two, 68.9% to 86.7% of outputs scored 4 or 5 of 5, by site and task. |
| Sun, Chen and Magrabi, 2018 | A pilot of a voice-activated, five-question interface for digital health incidents | It was usable. Participants worried about speaking aloud about sensitive safety issues in busy clinical areas. |
None of these tested whether chat changes what staff report. An older study frames the problem. In 2012 the HHS Office of Inspector General found hospital systems captured an estimated 14 percent of patient harm events.
Administrators said staff did not see 61 percent of all events as reportable. A friendlier form does not change what staff think is worth reporting.
Put human review on every field that matters
The studies show why. Models over-apply labels and are weakest at contributing factors. In the obstetric study, nearly half the model's labels were not on the reviewer's list. In the Swiss study, contributing factors were the least accurate task, at 72%.
IncidentKit is built around that. Lauren, the AI assistant, asks the follow-up questions a risk manager would ask, fills the form and drafts the investigation. Every AI-drafted field shows "Lauren · draft" until a person approves it. A person always reviews, edits and signs.
The audit trail records every change.
Scrutinize most where an error costs most: causes, harm level, and any classification that starts a clock, such as whether an event is reportable to the state or recordable under OSHA.
Seven things to watch
Most are about trust: the reporter's trust in the process, and your trust in the draft.
- Over-labeling. Count how often reviewers remove tags the report does not support.
- Invented causes. A cause the reporter did not state should read as a question.
- Leading questions. "Was the alarm off?" shapes the answer. Keep wording neutral.
- Privacy. Patient information sent to an AI provider needs a BAA covering that provider.
- Speaking aloud. Voice in open clinical areas worried pilot participants. Offer text.
- Language access. Reporters need their own language. Spanish and others are rolling out.
- Fallback. If the assistant is down or unwanted, the web form must still work.
Ask which provider handles your data, how long it is kept, and whether it trains models. See HIPAA and security. In IncidentKit, text is live and voice is rolling out.
IncidentKit offers several doors, with Lauren behind one
Staff can describe what happened by text, scan a QR code for a quick report, email email-to-incident, or use the web form. Voice reporting is rolling out.
On the free Open plan, Lauren intake has a daily limit. On the per-site Regulated plan, Lauren runs on a BAA-covered AI provider and patient information is allowed. See Lauren and AI for incident reporting.
The studies above did not test IncidentKit, and we publish no accuracy numbers for it. Run your own pilot.
A plain form is still right in some cases
A form is a good tool. Nothing in the evidence says every organization needs a chat. Choose a form when:
- Volume and variety are low.
- Fixed fields are enough and the narrative adds little.
- A policy or contract bars AI processing of the data.
- Nobody has time to review drafts the same day.
- Staff prefer a form and report well with it.
Pilot AI intake in five steps
Start small, run it beside your current form, and measure what the literature does not.
- Pick one unit or siteTell staff what the assistant does, that a person reads every report, and that they can still use the form.
- Run it beside your formKeep both open for several weeks to compare the same kinds of events.
- Measure what mattersTime to submit, completeness, how often reviewers change AI-drafted fields, and report counts including near misses.
- Check drafts against the reporter's wordsSample weekly for over-labeling, invented causes and leading questions.
- Write the rule with staffWho reviews, how fast, what happens when the assistant is down, and who can turn it off.
Frequently asked questions
Does AI write the incident report?
No. In IncidentKit, Lauren asks follow-up questions and drafts the fields. Every AI-drafted field shows "Lauren · draft" until a person approves it. A person always reviews, edits and signs, and the audit trail records every change.
Will AI intake increase the number of reports?
We found no peer-reviewed study showing that it does. The case is a lower barrier: less typing, no blank form. But HHS OIG found staff did not see 61 percent of harm events as reportable, and a better form does not change that. Measure report counts in your own pilot.
Is it safe to put patient information into an AI incident tool?
Only with a business associate agreement that covers the AI provider. HIPAA requires satisfactory assurances in a written agreement. IncidentKit allows patient information on its Regulated plan, which includes a BAA and Lauren on a BAA-covered AI provider. The free Open plan is for non-patient incidents.
Can AI work out the root cause of an incident?
It can suggest contributing factors to consider, but that was the weakest task in the studies. In one study of 10,063 Swiss incident messages, a model extracted events with 92% accuracy but contributing factors with 72%. People who know the process should judge causes.
Does AI intake replace the risk manager?
No. It moves the risk manager's time from chasing missing details to reviewing drafts and judging causes, actions and regulatory duties. A person owns every decision that matters. Reviewing drafts takes time too, so plan for it.
Sources
- Johnson et al. (2024): Accuracy of a proprietary large language model in labeling obstetric incident reports
- Denecke and Paula (2025): Evaluating large language models for analysing safety risks in healthcare incident reports
- Dobranski et al. (2026): Automated analysis of radiation oncology incident reports using large language models
- Sun, Chen and Magrabi (2018): Voice-activated conversational interfaces for reporting patient safety incidents
- HHS OIG: Hospital Incident Reporting Systems Do Not Capture Most Patient Harm (OEI-06-09-00091, January 2012)
- 45 CFR 164.502: Uses and disclosures of protected health information (business associates)
Reviewed against the sources above on Oct 5, 2026. Rules change: confirm current requirements with the issuing body or your counsel before relying on any summary.
Start with one incident.
Create your kit in about ten minutes and report the first incident the same day. Free to start, no card.