What AI gets wrong and what failure teaches us
Oct 6, 2026, 9:19 AM · Microsoft Research

A Microsoft researcher's work points to the quiet failures of AI at work, models that drift across long conversations and slowly corrupt documents, which benchmarks rarely catch.
Why it matters
On the Microsoft Research Podcast, partner research manager Jennifer Neville, who also holds a chaired professorship at Purdue, described her team's work on how AI behaves in realistic, multi-step work. One finding: when a task that models handle well in a single, fully specified prompt is instead spread across several turns of clarification, performance degrades significantly.
Her team also studies long workflows like repeated document edits, where small uncaught errors pile up and leave a subtle loss of meaning in the work. Those are the failures that matter as AI moves from answering questions to doing jobs.
From the desk
We think this is one of the more useful conversations about AI reliability we've heard in a while, because it targets the gap between how models are tested and how people actually use them. Real users rarely spell everything out in the first message. They figure out what they want as they go. If models get worse exactly as people behave naturally, that's a design problem, not a user problem.
The document-corruption finding worries us more. A hallucinated legal citation is at least checkable. Content that slowly loses meaning across dozens of agent edits is not dramatic enough to notice and may never get flagged. If agents become the default way offices revise contracts, reports and code, this kind of drift could quietly degrade a lot of work before anyone measures it.
Neville's practical advice is sensible and honest: check the answers, don't treat these systems as ready for full delegation, and when a conversation goes off the rails, start over with one complete prompt. We'd add that the burden shouldn't stay with users forever. Her call for people to give detailed feedback when things go wrong makes sense for improving models, but it also asks users to do unpaid quality work. She says her team analyzes usage only at scale in privacy-preserving ways, without eyes on individual data. We take that at face value and would like to see it documented.
Our read: AI is useful at work today with a human in the loop, and research like this is how it gets more trustworthy. Her own hard-won lesson applies to everyone building agents: look at the data.
Context
Neville leads Microsoft Research's AI Interaction and Learning team, which designs evaluations around multi-turn, collaborative and long-horizon tasks and then uses the gaps it finds to improve models, including through reinforcement learning.
Who feels it
- Knowledge workers
- Restarting a confused conversation with one complete prompt is a practical fix, and checking outputs still matters.
- Agent builders
- Long-horizon workflows need checks for gradual drift, not just obvious wrong answers.
- Evaluation researchers
- Single-turn benchmarks overstate how well models perform in real, messy use.
What to watch
- Published results from Microsoft's reinforcement learning work on multi-turn reliability
- Industry benchmarks that measure gradual document degradation in agent workflows
- Agent products that add rollback or self-checking for long tasks
Companies: Microsoft