Why the Same AI Product Succeeds at One Company and Fails at Another
Every company that deploys an AI agent eventually hits the same wall. The dashboard says sessions are happening. It does not say if those sessions actually worked. A ten-turn conversation could mean a user got exactly what they needed and kept going out of enthusiasm, or it could mean they tried the same request ten different ways and left frustrated. Traditional monitoring cannot tell the two apart, and as more daily work quietly moves through AI agents, that blind spot gets more expensive every year.
Moritz Sudhof is the Co-founder & CEO of Bigspin AI, a company built to close that gap. Sudhof spent 15 years at the intersection of language, behavior, and machine learning, including a run as VP of AI at BetterUp, where he helped ship an AI coaching product used across hundreds of organizations. Bigspin AI reads every conversation between a company's users and its AI agent, not a sample of them, and turns that into a daily report on what is breaking down and why.
In this episode of Lead with AI, Dr. Tamara Nall speaks with Sudhof about the research behind Bigspin AI, the moment two identical AI products produced wildly different results for two different companies, and why he believes the biggest opportunity in AI right now has less to do with model intelligence and more to do with the conversation happening between a person and a machine.
The Metrics That Cannot Tell You What Actually Happened
Before Bigspin AI, Sudhof spent years building conversational AI at BetterUp, a coaching company, where his job was to ship a language model that could actually coach someone. He learned quickly that technical skill mattered far less than he expected. What mattered was understanding how people were actually talking to the AI coach, what they wanted from it, and what they needed from it, questions with very different answers than the same questions asked about a human coach.
Sudhof later collaborated with Stanford professor Chris Potts on a research study that found something counterintuitive: most of what determines if an AI interaction succeeds or fails is not decided by how the system is designed. It is decided in the actual back and forth between a person and the AI, at a level most product teams cannot see. A prompt can be well written. A model can score well on every benchmark available. Neither one guarantees the user on the other end leaves the conversation satisfied.
A Product Manager, a Data Scientist, and a Researcher That Never Sleep
Sudhof compares Bigspin AI to having a product manager, a data scientist, and a user experience researcher who read every conversation an AI agent has, every day, and report back on what they found. The hardest part, in his telling, is not gathering the data. It is understanding failure inside a conversation, which has no answer key. There is no error message when a user gets quietly more frustrated with each reply.
Bigspin AI is built to read those human signals, tone, hesitation, and pushback, at scale, then turn them into evidence a team can act on rather than a hunch a support ticket might confirm weeks later. Teams connect a telemetry source they are likely already using, such as LangSmith or Braintrust, and Bigspin AI runs analysis across every message, tool call, and trace rather than a sampled slice of them.
Same Product, Same Prompt, Two Completely Different Outcomes
One story from Sudhof's own customers shows exactly why that matters. A company deployed an AI agent to a few hundred users inside one organization, and it worked. Engagement was strong, feedback was positive, and the team felt ready to scale. They rolled the same agent out to a much larger organization, several thousand users this time, and the response was silence. Usage numbers looked fine on a dashboard. The excitement did not follow.
When the company brought in Bigspin AI to look at the actual conversations, the difference became clear. The new group of users had the same job titles as the first group, but they behaved nothing alike. They wrote shorter prompts. They pushed back less. In Sudhof's words, they were treating the agent more like a vending machine and less like a collaborator, dropping in a request and expecting a result rather than working with the tool the way the first group had. Once the company understood that mismatch, they adjusted how the agent responded to that kind of user, and results at the larger organization began to catch up to the first.
The Email That Mattered More Than the Model
The clearest example of this came from what Sudhof calls one of his biggest aha moments. Two user groups were using the exact same AI product inside the exact same company: same model, same prompt, same tools, same interface. It was a role play tool built to help people prepare for difficult conversations. One group showed roughly twice the improvement in preparedness that the other group did. On a standard dashboard, both groups looked identical. Both were using the product about the same amount.
Only when Sudhof's team read the actual transcripts did the difference appear. The more successful group was simply more open with the AI. They shared more of what they were actually afraid of and what they actually wanted out of the conversation. The less successful group held back. Tracing the difference further, the cause was not the AI at all. It was the introduction each group received before they ever opened the product. The wording of the email, and the framing on the welcome screen, had been different for each group. That framing changed how people showed up, and how people showed up changed what the AI could give them back.
The AI never changed. Only the introduction did.
The lesson stuck with Sudhof: some of the strongest levers for improving an AI product are not inside the AI system at all. They sit in the language and context surrounding it, the parts a team might never think to test.
Why Sudhof Wants AI Treated Like a Collaborator, Not a Vending Machine
Sudhof is direct about what worries him as AI takes on more consequential work, including tasks like helping flag disease risk or shaping decisions that affect someone's livelihood. His view is that anyone deploying AI has a responsibility to understand what it is actually doing once it is out in the world, beyond what it was designed to do on a whiteboard. Left unmonitored, he argues, an AI system can run for months while quietly misleading users or missing what they actually needed, with no one noticing until the damage shows up somewhere else.
That responsibility connects to the answer Sudhof gave when asked what humanity will regret not preparing for in this AI era. His answer centered on one distinction. A vending machine takes a request and returns a result, no engagement required. A collaborator pushes back, asks questions, and gets better the more a person engages with it. Sudhof believes the industry has leaned too hard toward the vending machine model, treating AI as something to delegate to rather than think alongside. The tools that deliver the most value, in his view, are the ones people are still willing to argue with.
What Happens When Most Conversations Are Between Two AI Agents
Asked where Bigspin AI and AI agents more broadly head by 2030, Sudhof describes a near future where each person runs roughly fifteen different agents handling dozens of background tasks, and where most conversations happening at any given moment are not between a human and an AI at all, but between one specialized agent and another. As that volume of AI generated text and decision making grows, he expects the bottleneck to shift. The challenge will not be building smarter models. It will be building the visibility layer that lets people understand what all of those agents are actually doing, so oversight and judgment do not get lost in the volume.
Quick Answers
What does Bigspin AI do? Bigspin AI reads every conversation between a company's users and its AI agent, not a sample, and delivers a daily report on where the experience is breaking down and why, so product teams do not have to guess.
How is Bigspin AI different from a standard AI monitoring dashboard? Standard dashboards track structured metrics such as session counts or completion rates. Bigspin AI analyzes the conversations themselves for human signals like frustration, hesitation, and pushback, the details a session count cannot capture.
What is an invisible failure in AI monitoring? It is a breakdown in an AI interaction that never shows up in standard metrics or evaluation suites because nothing technically errors out, even though the user did not get what they needed.
Does Bigspin AI replace human judgment? No. Sudhof describes Bigspin AI as closer to an analyst that surfaces patterns and evidence for a team to act on, not a system meant to make decisions without a person reviewing them.
Who is Moritz Sudhof? Moritz Sudhof is the Co-founder & CEO of Bigspin AI. He spent 15 years working across language, behavior, and machine learning, including time as VP of AI at BetterUp, before building Bigspin AI with research collaborator and Stanford professor Chris Potts.
For product and engineering teams who suspect their AI agent is performing differently than the dashboard suggests, Bigspin AI is live at bigspin.ai. For more conversations with the founders building the next generation of AI, subscribe to Lead with AI on your favorite podcast platform.
Follow or Subscribe to Lead with AI Podcast on your favorite platforms
Website: LeadwithAIPodcast.com | Apple Podcasts: Lead-with-AI | Spotify: Lead with AI | YouTube: @LeadwithAIPodcast | Facebook: Lead with AI | Instagram: @LeadwithAIpodcast | TikTok: @LeadwithAIpodcast | Twitter (X): @LeadwithAI
Follow Dr. Tamara Nall
LinkedIn: @TamaraNall | Website: TamaraNall.com | Email: Tamara@LeadwithAIPodcast.com
Follow Moritz Sudhof (Co-founder & CEO, Bigspin AI)
LinkedIn: @sudhof | Website: bigspin.ai | LinkedIn (Bigspin AI): @bigspin-ai

Comments