Turning a stalled AI feature into a tool auditors actually used.
Auditors found it hard to understand and check the changes an AI tool suggested, while stakeholders pushed for different redesigns.
auditors using the feature, beta vs. after launch31 → 47%Among auditors with work to review · April–May 2026
The challenge
An AI feature suggested changes that auditors then had to review and approve. Technically it worked, but auditors found its output hard to understand and check, and stakeholders each had their own idea of how to redesign it. The team needed to decide what to fix before paying for another redesign.
What I owned
I led the product work from start to finish. I researched the problem with auditors, tested AI prototypes with them, decided what went into the release and coordinated delivery with engineering and data science.
The decisions
- Find out why auditors didn’t trust it. Research showed four reasons: the AI sometimes labelled changes wrongly, it didn’t show where information came from, it gave too much detail without a summary, and its results went out of date when new evidence arrived.
- Keep the AI’s suggestions apart from human edits. Auditors needed to see exactly what the AI had changed, while their own edits went into the document’s history.
- Prototype the review screen and get stakeholders to agree on what mattered most: short summaries, a clear view of what changed and a direct link to the evidence.
- Fix those four problems first. Other requests, like detailed scoring or approving each change one by one, could wait.
What changed
- The team shipped summaries, changes grouped by type, a before-and-after view of the text, links to sources and a way to refresh results on demand. Auditors still made the final call.
- Among auditors who had work to review, the share using the feature rose from 31% in beta (1 April–3 May 2026) to 47% after launch (4–31 May), 16 points higher.
- It reached 56% in the last week of May, and 44% of users came back to it within 14 days. These numbers come from the rollout itself; we didn’t run a controlled test to isolate the effect of the redesign.
Make AI output easy to understand and check before adding more features.