Designed an AI system that surfaces trends, anomalies, and comparisons the moment a chart is inserted in Excel. Increased chart retention by 15%, validated Copilot’s value for non-coders, and repositioned Excel as an analytical assistant—not just a calculation tool—with 65% positive sentiment.
Liked this project?
Let's talk about what we can build together.
Introduction
Excel rendered the chart and stepped back. It never said what the chart meant. About half of inserted charts got deleted in the same session, and among Copilot-enabled users, roughly 17% were already charting versus about 8% who weren't. Only 1 in 10 used Copilot regularly. I was Lead Designer for Copilot Chart Intelligence at Microsoft. We shipped an auto-opening insights panel to an Insider ring of ~5,000 users: 65% thumbs-up against a 50% goal, >95% factual accuracy on 500+ reviewed insights, 40% panel interaction. Then two problems forced a redesign: unpredictable generation time, and a head-on collision with a second AI feature over the same 300 pixels. I proposed the fix and both teams adopted it. It shipped too, to our own 2% dogfood ring, about 1,000 users, before a re-org halted it short of the wider release V1 got.
Results Overview
Numbers first, before the story.
| Signal | Result | Target |
|---|---|---|
| Positive feedback (thumbs up) | 65% | 50% |
| Factual accuracy, manual review of 500+ insights | >95% | >90% |
| Panel interaction (hover, scroll, or click) | 40% of the treatment group | 20% |
| Median dwell time | 8 seconds | >5 seconds |
Panel Interaction
Positive Feedback Rate
Median Dwell Time
Factual Accuracy
The Problem: Charts Without Context Are Decoration
Excel renders numbers into bars and lines instantly. It never said why any of it mattered. Users made a chart, then spent the next stretch of the afternoon translating it into bullet points by hand for a manager who was going to ask "so what?" anyway.
About half of inserted charts got deleted in the same session. The measured range runs 45-58% depending on the day, settling near 47% in production. Competitors had already moved: Tableau's "Explain Data," Gemini in Sheets surfacing declining categories, ChatGPT Data Analyst writing plain-language summaries of uploaded files. Excel, the most-used data tool in the world, still rendered the chart and stepped back.
The business case: among Copilot-enabled users, roughly 17% created charts versus about 8% of non-enabled users. Charting was already Copilot-adjacent behavior, roughly double. But only 10.8% of enabled users used Copilot regularly. That six-point gap was the opening: people already doing analytical work, with an assistant they weren't touching.
Users weren't struggling to make charts. They were struggling to extract meaning from them. Excel gave them the "what." It never gave them the "so what." That's not a charting problem. It's the reason a chart, once made, gets thrown away.
What Users Were Telling Us
Verbatim signals from research and OCV — the qualitative layer that made the quantitative data undeniable.
"The chart shows the data, but I still don't know what it means."
— NPS Feedback
"I spend more time writing bullet points about the chart than creating it."
— Enterprise User
"My manager asks 'so what?' and I have to explain manually."
— Analyst
Key Observations
Patterns across all research tracks — the moments where the data started telling a consistent story.
“Charts should tell a story – ours don’t.” This succinct user insight underscored how Excel charts lacked narrative value
Many users deleted charts soon after insertion, signalling they didn’t find them useful or attractive - about half of inserted charts get deleted in the same session (45-58%, settling near 47% in production).
On Excel Web, charting drew a disproportionate share of negative feedback (22% of all Excel Web frown feedback was chart-related), citing missing features and difficulty “getting charts to tell me anything”
CORE INSIGHT: Users weren't struggling to make charts — they were struggling to extract meaning from them. Excel gave them the 'what' but never the 'so what?'
Why Competitors Were Winning
30+ tools analyzed to understand the intelligence gap Excel hadn't closed — and where the whitespace was.
In our competitive analysis, tools like Tableau, Power BI, and AI-native platforms were delivering automatic insights:
Tableau: 'Explain Data' feature surfaced statistical anomalies
Google Sheets + Gemini: Proactive suggestions like 'This category is declining'
ChatGPT Data Analyst: Generated natural language summaries of uploaded CSVs
Napkin AI: Auto-generated visual explainers from text descriptions
CRITICAL GAP: Competitors understood that modern data viz isn't just rendering pixels — it's helping humans think. Excel was stuck in the old paradigm.
The Business Case for AI Insights
Not just a user request — a strategic imperative tied directly to Copilot adoption and platform retention.
This wasn't just feature parity — it was strategic necessity:
The Opportunity
The gap between what Excel showed users and what they actually needed to understand their data.
Why Chart Insights, Why Now: These competitive insights underscored that to stay relevant and delight users, Excel had to infuse intelligence directly into charting. It wasn’t enough to improve the UI or add new chart types; the next logical step was a Copilot-driven experience where the moment a user creates a chart, the software adds value by explaining the data.
| Metric | Value | Denominator | Meaning |
| 2% | 8M / 400M | Total Excel MAU | Overall market penetration |
| 16.5% | 4M / 24M | Copilot-enabled users | Chart usage among Copilot users |
| ~11% | 2.6M / 24M | Copilot-enabled users | Copilot usage among enabled users |
Among Copilot-enabled users, roughly 17% created charts versus about 8% of non-enabled users, roughly double. Only 10.8% of enabled users used Copilot regularly. That six-point gap was the opportunity: people already doing analytical work, with an assistant they weren’t touching.
Competitors were turning charts into “visual narratives” – combining charts with insights and even action suggestions
JTBDs
Understanding what users are actually trying to accomplish when they create and analyze charts in Excel.
| When I... | Insert a chart to visualize quarterly revenue data across product lines |
| I want to... | Immediately understand what patterns, trends, and anomalies exist without manually calculating statistics or staring at the chart for 10 minutes |
| So I can... | Confidently present findings to my manager, make data-driven decisions faster, and avoid missing critical business insights |
| Without... | Spending 20+ minutes manually analyzing every data point, second-guessing my interpretation, or relying on my manager to spot issues I missed |
| When I... | Need to present analysis to stakeholders who don't have time to dig into raw data |
| I want to... | Have ready-made narrative bullets that explain what the chart shows in plain language, with the option to copy them directly into emails or presentations |
| So I can... | Save hours writing explanatory text, ensure I'm communicating the most important findings, and look like a data expert even if I'm not |
| Without... | Spending 30+ minutes writing bullet points, worrying I'm focusing on the wrong metrics, or having my manager ask 'what about X?' that I completely missed |
| When I... | Create a chart for a high-stakes presentation (board meeting, executive review) |
| I want to... | Get a second opinion from AI to confirm my interpretation is correct, or alert me to patterns I might have missed |
| So I can... | Present with confidence, avoid embarrassing mistakes, and discover insights that make me look smart rather than missing obvious trends |
| Without... | Asking my manager to double-check every chart, staying up late rechecking numbers, or getting called out in a meeting for missing something obvious |
| When I... | Work with data regularly but don't have formal analytics training |
| I want to... | See examples of how experienced analysts interpret data, so I can learn what questions to ask and what patterns matter |
| So I can... | Develop my analytical skills over time, become less dependent on others, and eventually spot these patterns myself |
| Without... | Taking a formal data analytics course, bothering my analyst colleagues with basic questions, or relying on trial-and-error that wastes time |
My Role: Designing Intelligence, Not Just Interfaces
I was Lead Designer for Copilot Chart Intelligence at Microsoft, responsible for:
MVP — Prove the Core Value
Show static insights on chart insertion to validate whether automatic, immediate analysis delivers value.
| Success criteria | 40% click rate achieved vs. 20% target |
| Validation | A/B test for chart retention, Copilot activation |
| Performance target | P95 <20s generation |
| Interaction | Click → popover → thumbs feedback → Ask Copilot |
| Constraints | Creator-only, no persistence, native charts only |
| Scope | Auto-opening insights panel on chart insert (Web first), 1-3 insights, manual refresh |
Make It Interactive & Contextual
Enrich insights with responsiveness to chart edits and expand access to chart consumers.
| Responsive insights | Auto-update when chart type/data changes |
| On-demand access | Right-click any chart → Generate Insights |
| Persistent insights | Save as chart property, visible to collaborators |
| Enhanced interaction | Copy text, "Explain why" button, hide individual insights |
| Platform parity | Win32, Mac support, multi-chart scenarios |
| Edge cases: | Trendlines, empty charts, error recovery |
Deep AI Analysis & Storytelling
Transform insights into a conversational analytical assistant with cross-chart narratives and M365 integration.
| Conversational analysis | Embedded Copilot chat with context maintenance |
| Cross-chart insights | Multi-chart narratives and high-level summaries |
| Visual highlights | Link insight text to chart elements (hover to highlight) |
| M365 integration | Send to PowerPoint with auto-generated slides |
| Advanced analytics | Predictive trends, correlations, diagnostic analysis, benchmarking |
| Vision | Excel as AI-driven analysis platform with intelligent partnership |
The MVP : Version 1
If Copilot surfaces insights the moment a chart is inserted, users will keep the chart, understand their data faster, and trust Excel as an analysis tool, not just a calculation tool.
The solution: when you insert a chart, Copilot immediately answers "What does this show?" Plain language. Specific. Something like "North region is the top contributor with 35% of total sales." Right next to the chart, before you've had to think about it.
Before any of the interaction design, we ran a moderated study on placement: six participants, two prototypes. One put the insights callout next to the chart. The other put the same content in the Copilot side pane. The callout won, and the reasoning was consistent: it stayed visually attached to the thing it explained. One participant put it more bluntly than any usability report usually gets: "The pane is overwhelming and I would immediately exit by looking at that."
I'll say the part I'd rather not: I have three summary artifacts from that study and they don't agree on the exact split. One says 4-of-6 preferred the callout, another says 3-of-6, and two participants swap sides depending on which document you read. Six is small enough that a two-person swing is most of the result, so I don't quote a count. What didn't move across any version of the write-up was the reasoning: proximity to the chart beat completeness of the panel, for every participant, including the ones who ended up preferring the pane.
Two candidates: a first-run opt-in gate, or auto-open with an easy dismiss and a settings toggle to switch it off. I argued for auto-open, and shipped it, for a reason that has nothing to do with making the feature look impressive: an opt-in gate only measures who says yes to a description of a feature. Auto-open with a dismiss and a toggle produces three signals a gate can't: dismiss rate, permanent-off rate, and re-enable rate. That third one only exists if the default starts on. Dismissal wasn't a failure metric I hoped would stay low. It was the point. We pre-registered five signals before launch: dismissal, dwell, copy, refresh, feedback. The honest cost: it's an intrusive default on a surface where people are mid-task in a financial model. I'd defend it as a controlled-ring learning configuration, not as a permanent GA default.
If dismissed users can trigger back insights from right-click menu or through chart ribbon menu
Another version exploring the Copilot pane instead of on-canvas dialog
The Collision: When Reality Punched Back
The first integration build returned insights in over twenty seconds against a five-second design assumption. A dogfood user said it plainly: "I inserted the chart and just waited. It felt broken. I thought Excel crashed." Engineering brought that down, but generation time stayed variable, and variable is the property that matters: an auto-opening panel can't promise when it will fill.
Independently, mid-sprint, we discovered Chart Insights and a second team's AI Design Recommendations feature triggered on the same "Insert Chart" event and wanted the same 300px of pane. On the 1366x768 laptops much of the commercial base runs, two panes plus the ribbon leave under 650px of actual workbook. Neither team had mapped shared trigger events. We found out in week six.
By ship, the pipeline returned insights in under two seconds for typical datasets and under five for larger ones, in 95% of cases. Engineering did bring the number down. What didn't go away was variability, because of why it's slow: the pipeline isn't one model call. It reads the schema, generates Python to compute the actual statistics, executes that code, then hands the summary back to the model to write the sentence. Four stages, one of them code execution, over a dataset whose size the user chooses. That's also the reason the insights don't hallucinate numbers: the model never computes anything, it only narrates what Python already calculated. The accuracy and the unpredictable timing come from the same architectural choice.
We designed for <5s latency, got >30s reality. Next time: prototype with artificial delays from day one. Assume performance will be worse than promised.
I inserted the chart and just... waited. It felt broken. I thought Excel crashed.
Users clearly preferred on-canvas overlay, but technical and real estate constraints made it impossible in its original form. The skittle button preserved the contextual proximity users loved while solving latency and space problems.
Version 2: Introducing the Skittle
I proposed a small on-chart control, the Skittle. It answered both problems with one design: ambient and non-blocking to solve the latency, zero permanent canvas cost to solve the collision. It's adoptable by both teams for the same reason it's correct for either one alone.
The other team adopted it as the shared pattern. It shipped too, to our own 2% dogfood ring, about 1,000 users on Excel Web, engineering-complete rather than a partial build. It never reached fastfood: a September 2025 re-org dissolved the team and reassigned priorities before it could ramp further.
I don't have a telemetry export for that ring, so what I can tell you is recollection, stated as ranges rather than points: roughly 80% of chart-insert sessions saw the control (75-85%), about a quarter of those who saw it tried it (20-30%), and dismissal ran low, around 7% (5-10%). That's the honest shape of it: shipped, dogfood-validated by memory, stopped by a re-org. I'd rather say that than overclaim a result I can't source, or underclaim that nothing came of the redesign at all.
We also added a pop-over toast to highlight
Key Design Decisions & Trade-offs
Thumbs, not stars; the callout, not the pane. Binary feedback over a five-point scale, because at this stage volume of signal mattered more than granularity. I needed to know which insights were wrong, not that one was a three versus a four.
The insight itself lives in a callout anchored to the chart, not the Copilot sidebar. A placement study (above) showed people wanted the explanation physically attached to the thing it explained. One participant said outright she'd exit on sight if it were a pane.
Choice: Auto-trigger on insert (with easy dismiss)
Why: Testing showed users didn't know to ASK for insights. Making it proactive was key to discovery.
Choice: No auto-refresh on data changes
Why: Performance risk, user distraction, and testing showed users preferred control. They'll hit refresh when ready.
Choice: Floating panel anchored to chart
Why: Sidebars compete with other UI. Inline feels contextual, less like 'another feature' and more like 'the chart explaining itself'
Choice: Binary thumbs up/down, not 5-star scale
Why: Lower friction, higher response rate. We cared more about volume of feedback than granularity.
Impact & Results
We validated this on Excel Web, with dogfood users and an Insider ring of roughly 5,000 users, against five pre-registered signals: dismissal, dwell, copy, refresh, feedback. We named all five before launch, which is the most defensible claim in this feature's record, because nothing was cherry-picked after the fact.
| Signal | Result | Target |
|---|---|---|
| Positive feedback (thumbs up) | 65% | 50% |
| Factual accuracy, manual review of 500+ insights | >95% | >90% |
| Panel interaction (hover, scroll, or click) | 40% of the treatment group | 20% |
| Median dwell time | 8 seconds | >5 seconds |
Two honest caveats belong next to those numbers, not buried in a footnote. The 40% is broad interaction: hover and scroll count, not just clicks on a control. And people rated individual insight statements helpful about 55% of the time, with roughly 10% marked not useful. That's a more sober number than the 65% panel-level score, and it's the one I'd work from if I were deciding what to fix next.
What's not in this table: I don't have a reconciled figure for chart retention or a "2x more likely to be kept" claim. Neither appears anywhere in the evidence corpus I reconciled against. Not cleared, not barred. Just absent. I'm flagging this rather than restating it: either a source exists that this pass didn't locate, or the figure needs to come down until one does.
V2, the on-chart control from the redesign, was fully specified. The Chart Insights and Design Recommendations tracks both adopted it as the shared direction, and it shipped: to our own 2% dogfood ring on Excel Web, about 1,000 users, engineering-complete rather than a partial build. It never reached fastfood. A September 2025 re-org dissolved the team and reassigned priorities before it could ramp further.
I don't have a telemetry export for that ring. What follows is recollection, given deliberately as ranges rather than points:
| V2 signal (dogfood ring, recollection tier) | Result |
|---|---|
| Seen, of chart-insert sessions | ~80% (range 75-85%) |
| Tried, of those who saw it | ~25% (range 20-30%) |
| Dismissed | ~7% (range 5-10%) |
Two things worth saying about that table before anyone else has to point them out. First, an unrelated on-cell Copilot control on Win32 also reports "~80%." But that one is a relative increase in seen rate on someone else's surface, not an absolute rate on ours. Same digits, different metric, coincidence, not corroboration. I say "increase" for theirs and "of sessions" for ours every time so the two don't collide in a reader's head. Second, there's no clean line from these numbers to V1's 40% interaction figure above: different container, a fifth the ring size, and a few weeks of dogfood instead of a full fastfood window before the halt. I wouldn't put V1's and V2's numbers in the same sentence without saying that.
Critical feedback we addressed:
Key Learnings
What three design iterations, a latency crisis, and a feature collision taught me about shipping AI features that earn trust.
Built for <5s latency, got >30s reality. Had to pivot mid-sprint.
Next time: Design for 10x slower than best case. Prototype with artificial delays (5s, 15s, 30s). Have backup pattern ready from day one.
Designed in isolation, learned about Design Recs conflict late.
Next time: Audit all features that trigger on same event. Test on 1366x768 screens from day one. Involve PM in multi-feature roadmap alignment earlier.
Defined metrics after design was done. Had to retrofit event tracking.
Next time: Create telemetry schema during wireframing. Every interaction state = logged event. Treat telemetry as a design deliverable.
Tested with clean datasets (5 columns, 100 rows), shipped to messy reality (500 columns, formulas, merged cells).
Next time: Start with messiest data first. Create 'data chaos test suite' for prompt validation. Design failure states as prominently as success states.
The Bigger Lesson
Map shared trigger events during discovery. Any two features firing on the same event are on a collision course. That's a one-hour audit in week one. It cost us a sprint in week six. It's the cheapest lesson in this project, and the one I'd take to any team, not just the one I happened to collide with.
Prototype latency variations earlier. Design for 10x slower than best case. Have the backup interaction pattern in the design file before the first engineering conversation.
Build instrumentation into design from day one. I defined metrics after the design was done. That meant retrofitting event tracking. Every interaction state should have a corresponding logged event. Treat the telemetry schema as a design deliverable.
Test with messy data from day one. We tested with clean datasets: five columns, 100 rows. We shipped to 500 columns, formulas, merged cells. Start with the messiest possible data. Build a failure state design that's as prominent as the success state.
The Newsletter
Conclusion
AI features don't earn trust by being impressive. They earn it by being right.
Chart Insights worked because it solved a specific job someone actually had (understand my chart right now), appeared at the exact moment they needed it (chart insertion), told the truth more than 95% of the time, and left the person in charge of the analysis. The AI helped them see faster. It didn't replace what they were doing.
That's the model. Not AI for its own sake. AI that has a clear job and does it reliably.