Skip to main content
General

When One Agency Replaced Its Note-Taker With Software

A

Agency Script Editorial

Editorial Team

August 27, 2017·8 min read
ai note-taking and summarization appsai note-taking and summarization apps case studyai note-taking and summarization apps guideai tools

This is the story of a mid-sized creative agency that spent a quarter moving from human-typed meeting notes to an AI-driven workflow. The details are composited from common patterns, but the arc is real: the situation that forced a decision, the choice they made, how the rollout actually went, and what the numbers showed at the end.

It is worth telling because the agency did not succeed on the first try. Their initial deployment produced a small disaster, they nearly abandoned the tool, and then a few targeted changes turned it into one of the most valued parts of their operating rhythm. The lessons are in the recovery, not the launch.

The Situation

The agency ran a high volume of client meetings, and notes were a constant source of friction.

The Pain

Account managers split attention between listening and typing during client calls, then spent up to an hour afterward writing recap emails. Notes were inconsistent across people, action items got lost between the call and the project tool, and when a client disputed what had been agreed, there was no reliable record to point to. The cost was both hours and trust.

The Trigger

Two missed deliverables in one month, both traceable to action items that never made it out of someone's notebook, pushed leadership to act. The mandate was simple: stop losing commitments and stop burning hours on recap writing.

The Decision

They chose an AI note-taker that could join calls, produce structured summaries, and integrate with their project management tool.

The selection criteria were deliberate. They needed solid integration with their existing stack, configurable retention, and a custom vocabulary for client and project names. Price mattered less than reliability and data terms, since these were client conversations. The reasoning behind those criteria mirrors what we lay out in Choosing Among Otter, Fathom, and the Summarizer Crowd.

What they got right at this stage was naming the primary job before shopping. The job was follow-through and reliable client recaps, not building a searchable archive. That clarity let them weight integration and accuracy over breadth of features, and it kept them from being seduced by a tool that did many things adequately rather than the few things they needed well. What they got wrong came later, and it had nothing to do with the tool they picked.

The First Rollout and Why It Failed

They turned it on across all client calls, told the team it was live, and stepped back. Within three weeks, problems surfaced.

What Went Wrong

A recap email sent to a client listed a commitment the agency had never made, because the model summarized a hypothetical discussion as a decision. Client names were misspelled throughout because nobody had configured vocabulary. And one summary attributed a budget concern to the wrong stakeholder, creating an awkward follow-up. Trust in the tool cratered, and several account managers quietly went back to manual notes.

The mistakes here were textbook, and they map directly to the patterns in Where Meeting Notes Quietly Go Wrong With AI Transcription.

The Recovery

Instead of abandoning the tool, leadership treated the failures as configuration problems and process gaps.

The Changes They Made

They loaded a custom vocabulary of every client and project name. They added a hard rule that no client-facing summary went out without the account lead reviewing and correcting it against the transcript. They wired extracted action items directly into the project tool with owners and due dates. And they set retention windows so transcripts expired on a schedule rather than accumulating. These changes echo the practices in Habits That Keep AI Meeting Notes Trustworthy.

The Execution and Outcome

The second rollout was incremental. They piloted with two account teams for a month, refined the review workflow, then expanded. The incremental approach mattered because it let them catch the remaining rough edges with two teams rather than twelve.

The Measurable Results

Recap email time dropped from roughly an hour to about fifteen minutes per call, since account managers were now editing a verified draft rather than writing from scratch. Lost action items, the original trigger, went to near zero once items flowed into the project tool automatically. Client disputes became easier to resolve because a searchable transcript existed. The human-review step added a few minutes per client summary, which leadership considered a clear trade for the reliability it bought.

Leadership was careful about how it measured this. They deliberately ignored the headline number the tool reported, the count of meetings summarized, because it said nothing about value. Instead they tracked the two things tied to the original problem: recap time per call and the rate of action items that actually got completed. Both moved in the right direction within the pilot month, which gave them the confidence to expand. Anchoring the metrics to the business problem, rather than to tool activity, was what made the expansion decision defensible.

The Cost Side

It was not free. The review discipline required ongoing attention, and the team had to resist the temptation to let unverified summaries slip out when busy. But measured against two missed deliverables a month, the math was not close.

There was also a softer benefit that leadership had not predicted. Account managers reported being more present on calls, because they were no longer splitting attention between listening and typing. Clients noticed the difference, and a few commented that meetings felt more focused. None of this showed up in the original business case, which was about lost commitments and recap hours, but it became one of the reasons the team defended the tool when budget season came around.

The Lessons

The agency's experience distills into a few transferable lessons.

First, defaults are decisions, and turning the tool on everywhere with no configuration is itself a choice with consequences. Second, the value is unlocked by routing and verification, not by the raw summary. Third, a failed first rollout is recoverable when you treat the failure as a process gap rather than a verdict on the technology. The way they tracked improvement is detailed in Numbers That Tell You an AI Summarizer Is Working.

A fourth lesson is about sequencing. The agency's mistake was not adopting AI notes; it was adopting them everywhere at once, with no pilot, no configuration, and no verification. Had they piloted from the start, the invented commitment would have surfaced with one internal team instead of in front of a client. The recovery essentially reran the adoption the way it should have gone the first time: small scope, configured tool, human checkpoint, then expansion. The technology was ready before the process was, and closing that gap was the entire story.

Frequently Asked Questions

Why did the first rollout fail?

They deployed with default settings and no verification step, so the tool invented a commitment, misspelled client names, and misattributed a concern. The technology was capable; the workflow around it was missing.

What single change had the biggest impact?

The mandatory human review of client-facing summaries against the transcript. It eliminated the embarrassing errors that had destroyed trust, at a cost of only a few minutes per summary.

How long did it take to see results?

The pilot ran one month with two teams before expansion. Meaningful time savings and the drop in lost action items were visible within that first month once the configuration and review steps were in place.

Was the time savings worth the added review step?

Yes, by a wide margin. Recap time fell from about an hour to fifteen minutes per call, while the review added only a few minutes. The net was a large reduction in effort plus far higher reliability.

Should other agencies copy this exact setup?

The principles transfer; the specifics depend on your stack. Configure vocabulary, route action items into your real tools, verify client-facing output, and set retention limits. The tool you choose matters less than getting those four things right.

What would have happened if they had abandoned the tool?

They would have returned to the original problem: lost commitments and an hour per recap. The failure was fixable, and treating it as a verdict on AI rather than a process gap would have left real value on the table.

Key Takeaways

  • Missed deliverables traceable to lost action items forced the move from manual to AI-assisted notes.
  • The first rollout failed because of default settings and no verification, producing invented commitments and misattribution.
  • Recovery came from vocabulary configuration, mandatory review of client summaries, action-item routing, and retention limits.
  • Recap email time fell from about an hour to fifteen minutes, and lost action items dropped to near zero.
  • The human-review step cost a few minutes per summary and was an easy trade for reliability.
  • A failed first rollout is recoverable when treated as a process gap rather than a verdict on the technology.
A

Agency Script Editorial

Editorial Team

The Agency Script editorial team delivers operational insights on AI delivery, certification, and governance for modern agency operators.

Ready to certify your AI capability?

Join the professionals building governed, repeatable AI delivery systems.

Explore Certification