This is the story of a mid-sized SaaS company's finance team and the year they spent replacing a forecast nobody trusted. The details are composite, drawn from patterns common to companies at this stage, but the arc is faithful to how these projects actually unfold: messier than the vendor demo, slower than the plan, and ultimately worth it for reasons the team did not anticipate at the start.
The company had roughly forty million in annual recurring revenue, a finance team of five, and a forecasting process built entirely in a sprawling spreadsheet. The spreadsheet worked, in the sense that it produced a number every month. It did not work in the sense that anyone believed the number.
The Situation: A Forecast on Life Support
The spreadsheet had grown for four years. It carried the assumptions of three different analysts who had since left. Each month, closing it took two full days, and the output was a single revenue figure with no sense of its uncertainty.
The breaking point
When the board asked for a downside scenario, the team could not produce one without rebuilding half the model by hand. The CFO realized the forecast could answer "what is the number" but never "how confident are we and what moves it." That gap is exactly what the practices in Disciplines That Keep an AI Forecast Trustworthy are meant to close.
The Decision: Buy Capability, Not Magic
The team evaluated several AI financial forecasting tools. They resisted the temptation to choose on algorithm sophistication and instead chose on three criteria: integration with their billing system, the ability to model drivers rather than aggregate revenue, and a native prediction interval.
Why those criteria
They had been burned by the spreadsheet's opacity, so explainability outranked raw accuracy. A model they could not interrogate would just be a faster way to distrust a number. The selection logic followed the reasoning in Sorting Through the Crowded Forecasting Software Market.
The Execution: Slower Than Planned
The rollout was budgeted at six weeks. It took eleven.
Where the time went
The data was the problem, as it almost always is. Billing data had inconsistent product names, churned accounts that were not flagged cleanly, and a pandemic-era spike nobody had documented. Before the tool could forecast anything useful, the team spent three weeks building an anomaly register and reconciling the billing export.
The first honest forecast
Once the data was clean, the team modeled new bookings, expansion, and churn separately. The tool combined them into an MRR forecast with an interval. The first output was sobering: the range was wider than leadership expected, which was uncomfortable and also the first honest thing the forecast had ever said.
The Validation: Earning Trust Through Backtesting
Rather than launch the new forecast into the board deck, the team ran it in parallel with the spreadsheet for two quarters and backtested both against actuals.
What the backtest showed
The AI-driven forecast was not dramatically more accurate on the central number. Its real advantage was that its prediction interval was honest: actuals landed inside the stated range as often as the math said they should. The spreadsheet, by contrast, had been quietly overconfident for years. The metrics they used are detailed in Reading Whether Your Forecast Engine Is Actually Working.
The Outcome: Measurable and Cultural
The numbers improved. Monthly close on the forecast dropped from two days to under three hours. The team could produce a downside scenario in minutes by adjusting churn and conversion drivers.
The change that mattered more
The cultural shift outweighed the time savings. Board conversations moved from arguing about a single point to discussing where reality was likely to fall and which drivers to watch. The forecast stopped being a thing the CFO defended and became a thing the leadership team reasoned with. The failure modes the team learned to avoid along the way are catalogued in Seven Ways Forecasting Models Quietly Mislead Finance Teams.
The Lessons Worth Keeping
First, the data work dwarfs the modeling work, and any timeline that assumes otherwise will slip. Second, an honest wide range beats a confident narrow one, even though it feels worse in the room. Third, parallel-running before cutover bought the trust that made adoption stick. Skipping that step would have left the team with a better tool nobody believed.
What the Team Got Wrong Along the Way
The story is not a straight line of good decisions. Early on, the CFO wanted to cut the parallel-run short after a single strong quarter, eager to retire the spreadsheet. The forecast owner pushed back, arguing that one quarter was not enough evidence to trust the new system in front of the board. That argument nearly became a conflict, and it was the right one to have. A single good quarter can happen by luck; two quarters of honest calibration cannot.
The team also initially tried to model consolidated revenue directly before realizing the drivers told a clearer story. They spent two weeks on an aggregate approach that produced a number nobody could explain, then scrapped it and rebuilt around bookings, expansion, and churn. That detour cost time but taught the team why explainability had to come first, a lesson that echoes the failures in Seven Ways Forecasting Models Quietly Mislead Finance Teams.
The detour that paid off
The aggregate detour, frustrating as it was, became the team's strongest argument internally. Having seen both approaches side by side, nobody questioned the move to driver-based modeling again. Sometimes the fastest way to win an organization over to the right structure is to let it watch the wrong one fail cheaply first.
How the Forecast Looks a Year Later
A year after cutover, the forecast had become invisible in the best sense. It ran each cycle, updated as billing data flowed in, and produced a range the leadership team reasoned with rather than argued about. The monthly error review caught one episode of drift early, when a pricing change shifted expansion patterns, and the owner re-decomposed the expansion driver before the drift reached a board number. That single catch, the kind of routine save the new process made possible, would have been a quarter-end surprise under the old spreadsheet. The tool did not make the team smarter. It gave their existing judgment a faster, more honest instrument to work with.
What Made This Project Succeed Where Others Stall
It is worth being explicit about why this rollout worked, because many do not. Plenty of finance teams buy a capable tool, configure it over a rushed few weeks, and quietly abandon it within a year because nobody trusts the output. This team avoided that fate for three reasons that had nothing to do with the software.
First, they chose on the right criteria, prioritizing integration and explainability over algorithm sophistication, so the tool fit the way finance actually worked. Second, they budgeted honestly for the data cleanup that always dominates these projects, which meant the inevitable overrun did not derail the effort or burn out the team. Third, they earned trust through evidence rather than asserting it, running in parallel and backtesting until the new forecast had a track record nobody could dismiss.
The transferable lesson
None of these three reasons is specific to this company. Any finance team can choose on integration and explainability, budget for data work, and earn trust through parallel-running. The reason most rollouts stall is not that they picked the wrong tool but that they skipped one of these three disciplines, usually the patient trust-building, because it feels slow. This team's willingness to be slow where it mattered is precisely what let the forecast become a durable part of how they plan, an outcome the broader practices in Disciplines That Keep an AI Forecast Trustworthy are designed to reproduce.
Frequently Asked Questions
How long did the whole project take?
Roughly a quarter from decision to full cutover, with most of the overrun coming from data cleanup rather than tool configuration.
Was the AI forecast much more accurate?
Not on the central number. Its advantage was an honest prediction interval and the ability to run scenarios quickly, which mattered more than a marginal accuracy gain.
What was the single biggest obstacle?
Dirty billing data, including an undocumented pandemic-era spike. Building the anomaly register and reconciling the export consumed most of the schedule.
Why run in parallel for two quarters?
To build trust through evidence. Backtesting both forecasts against actuals proved the new one was honestly calibrated before anyone bet a board decision on it.
What would the team do differently?
Budget for the data work realistically from the start and set expectations that the first honest forecast would show a wider range than the old spreadsheet implied.
Does this generalize beyond SaaS?
The specifics of the drivers change, but the arc holds for most finance teams: the data work dominates, calibration beats false precision, and parallel-running earns adoption.
Key Takeaways
- The team replaced a spreadsheet that produced a number nobody believed with a driver-based forecast leadership could reason with.
- They chose tooling on integration, driver modeling, and honest intervals rather than algorithm sophistication.
- Data cleanup, not modeling, consumed most of the timeline and caused the overrun.
- Parallel-running and backtesting for two quarters earned the trust that made adoption stick.
- The lasting win was cultural: conversations shifted from defending a point to reasoning about a range and its drivers.