Designed a Copilot feature that analyzes data shape, infers user intent, and recommends optimal chart configs with story-first titles like “Quarterly Trends.” Collaborated with ML engineers on RAG model tuning, eliminating chart type decision fatigue and making great visualizations accessible to all.
Liked this project?
Let's talk about what we can build together.
Introduction
How I designed an AI system that bridges the gap between chart creation and chart communication — for 400M Excel users.
Creating charts is easy. Making them GOOD is hard. Users struggled with chart type selection, styling decisions, and best practices—resulting in suboptimal visualisations even when data was correct.
As Lead Designer for Copilot Chart Design Recommendations, I designed an LLM-powered system that analyses data shape, infers user intent, and suggests optimal chart configurations. This required deep collaboration with ML engineers to train and tune the RAG models with data visualisation principles—essentially encoding expert knowledge into AI prompts.
The result bridges the gap between 'chart exists' and 'chart communicates effectively,' democratising data visualisation expertise for 400M users.
Results Overview
The feature shipped, scaled, and proved that AI-assisted design guidance moves real product needles.
Execution Success
User Effort Saved
Clicks Eliminated
Shipping Timeline
The Problem: The Chart Design Expertise Gap
Most users could create a chart. Almost none could create a good one. The expertise gap was the product gap.
"No suggestions... I have to try different charts and hope one communicates my idea." — Usability Study 2024
"I spend hours tweaking charts to look professional." — Power User
"It looks boring. How do I make it presentation-ready?" — Enterprise Analyst
"I created a chart but don't know if I picked the right type." — Intermediate User
| Pain Point | User Behaviour | Business Impact |
|---|---|---|
| Chart Type Uncertainty | Try multiple types, delete, start over. 5-10 minute cycle. | Wasted time, user frustration, suboptimal final choices |
| Styling Paralysis | Don't know which formatting options matter. Either over-style or under-style. | Charts look unprofessional or cluttered |
| Best Practice Ignorance | Unaware of data viz principles (e.g., start axis at zero, use direct labels) | Misleading visuals, poor communication |
The pattern was clear, Users could INSERT charts (thanks to our P0 improvements), but they couldn't optimise them.
Why Competitors Had the Advantage
Competitive analysis revealed sophisticated design assistance
Tools that handle complex data are hard to use. User-friendly tools handle simple data.
40% of Excel charts were deleted same session. Canva/Flourish users kept charts because they looked presentation-ready on first insert.
Google Explore, Napkin.ai, Tableau Show Me — every modern tool reduces data→chart to 1-2 clicks. Excel required 5+ steps with no guidance.
Pitch, Miro, Figma use click-to-format context menus. Excel used ribbon + dialog boxes — 3+ clicks to format a single element.
In our AI compete benchmark, Copilot in Excel scored 48/100 — below ChatGPT (85), Gemini (72), even Gemini Sheets (56). Task success was 40%.
BI tools provide insights alongside charts. Google Gemini explains trends. Excel charts were 'purely graphical — static visuals, with no story.’
Synthesis
Across all 30+ tools, three truths emerged:
Reduce friction at the start — suggestions, templates, one-click creation.
Make the default output impressive — users keep charts that look good on first insert.
Add intelligence — the tool should explain the data, not just display it. Excel had the data capability. It needed the ease and the storytelling.
As a User
Functional & Emotional JTBDs
What users wanted to accomplish — and how they wanted to feel — when reaching for charts in Excel.
"When I insert a chart, help me create a visualization that tells my story effectively — without needing to be a data viz expert."
| Job Category | User Statement | Pain Point It Solves |
|---|---|---|
| Chart Type Selection | "Help me figure out which chart best represents my data" | 5-10 minute trial-and-error cycles; users try multiple types, delete, start over |
| Visual Design | "Make my chart look professional and presentation-ready" | Charts described as "boring," "old-fashioned," "embarrassing to present" |
| Best Practice Application | "Tell me what I don't know about good data visualization" | Users unaware of principles like "start Y-axis at zero" or "use direct labels" |
| JTBD | Description | Copilot Intent Share |
|---|---|---|
| Comparative Analysis | Compare values across categories, geographies, or periods to uncover insights | Part of 83% "Create Chart" intents |
| Presentation & Storytelling | Make complex information clear, engaging, persuasive in meetings/reports | 9.6% of explicit intents |
| Trend Analysis | Visualize how metrics change over time to identify patterns | Primary use case for Line charts |
| Answering Business Questions | Create ad-hoc visuals to answer specific questions quickly | Core Excel workflow |
The "Magic Wand" Quote
from User Research
The single question that unlocked what users truly needed — and reframed the entire design brief.
"If you could wave a magic wand, what would you change?"
Users wanted three things:
Automatic chart creation
"Based on my specific goal and storytelling needs, help me tell my story"
Automatic beautification
"Make my charts look beautiful without me having to figure it out"
Natural language customization
"Let me ask for customizations in plain English"
As a Business
Strategic JTBDs
The commercial imperatives driving investment in chart intelligence — retention, ecosystem depth, and competitive parity.
"Increase chart adoption and retention to keep users within the M365 ecosystem for their data visualization needs — preventing defection to competitors."
| Metric | Baseline Problem | Target Impact |
|---|---|---|
| Chart Kept Rate | ~45% of charts deleted in same session | Push toward >70% retention |
| Chart Create MAU | Only 2% of MAU on web create charts | Increase top-of-funnel creation |
| Net Chart Creation | Inserts minus deletes was too low | Increase net positive |
| Data Viz NPS | Charting issues dragging down Excel NPS | Measurable improvement |
| Copilot Tried/Enabled | Design Recommendations as gateway | Lift adoption rate |
| Business Job | Why It Matters | How Design Recommendations Solves It |
|---|---|---|
| Compete Defence | Tableau, Power BI, ChatGPT Code Interpreter, Napkin AI democratizing design expertise | Embed expertise IN the tool; no learning curve required |
| Copilot Adoption | Only ~9% of Copilot users engaged with chart-related prompts | Proactive recommendations at insert = gateway to Copilot |
| User Retention | Users looking outside M365 for data viz needs | "Wow moment" on first chart = sticky behavior |
| Unlock Latent Demand | 33% of commercial users want to create charts but don't | Remove friction to convert intent → action |
The Business Funnel Problem
The drop from awareness to creation is where AI Design Recommendations lives — most people who know Excel can make a chart never actually create one.
My Role
Designing AI as Design Partner
I wasn't just designing a UI — I was co-designing the intelligence behind it, working across ML, data science, engineering, and research simultaneously.
Systems thinking — Thinking Charts through complete M365 ecosystem
RAG model training & tuning — defined data viz properties that inform chart type recommendations
Recommendation interaction patterns — preview, apply, undo flows
AI prompt engineering collaboration — co-designed LLM prompts with ML team for chart analysis, gave examples of visually stunning data viz.
End-to-end UX strategy for Copilot-powered design recommendations
Multi-recommendation handling — when LLM suggests 3-5 improvements, how to present without overwhelming
Trust-building mechanisms — explainability, rationale, learn more links
Design recommendations should feel like:
The design philosophy that shaped every interaction pattern, recommendation format, and piece of copy in the system.
A helpful colleague, not a know-it-all boss
Educational — explain WHY, don't just say WHAT
Suggestions, not mandates — users always have final say
Confidence-building — help users become better designers over time
Phase 1:
Analyzed telemetry revealing a steep funnel drop from chart awareness to creation. Synthesized OCV feedback — users called charts "boring" and "embarrassing." Defined the core JTBD: help users tell data stories without being viz experts.
OCV analysis pointed to poor chart quality — tables instead of charts, wrong grouping, blank outputs — as a recurring complaint. That's what we were solving.
Phase 2:
Explored 3 directions: auto-apply magic, inline tooltips, side-by-side preview. User testing rejected "AI takeover" — they wanted to see options first. Landed on story-first titles and preview-before-commit as guiding principles.
Phase 3:
Partnered with engineering to map hard limits: 2-4s LLM latency, ~85% preview fidelity, single Copilot pane. Made key tradeoffs — 4 recommendations, refresh button, dropdown for placement. Designed around constraints, not against them.
Phase 4:
Built a golden dataset of 50+ data scenarios with ideal chart recommendations. Defined statistical signals for preprocessing — time-series, part-to-whole, category comparison. Reviewed model outputs weekly to catch and correct bad patterns.
Phase 5:
Working with the ML team, I co-created prompts optimized for chart type selection. The key was encoding data visualization best practices into the prompt structure — story-first titles, rationale text, diverse recommendations.
I defined the decision tree that the model uses to recommend chart types:
| User Intent / Data Shape | Recommended Chart | Why This Works |
| Time series with trend | Line Chart | Shows change over time; eye follows the trajectory |
| Categorical comparison | Clustered Bar/Column | Easy side-by-side comparison; clear value differences |
| Part-to-whole (<7 categories) | Pie/Donut Chart | Intuitive percentage representation; limited categories |
| Part-to-whole (>7 categories) | Stacked Bar/Area | Handles many categories; shows composition |
| Correlation/distribution | Scatter/Bubble Chart | Reveals relationships; shows outliers clearly |
| Actual vs. Target | Combo Chart | Different visual encoding for different data types |
Through iterative testing on 20+ sample datasets across industries (Telecom, Finance, Manufacturing, Retail), we tuned the prompts to:
**Prompt Architecture**
Given a chart with [data structure], current type [X], analyze if a better visualization exists.
Consider:
1) Data relationships,
2) Storytelling intent,
3) Visual clarity
Return top 4 recommendations with executable chart config and brief rationale.Phase 6:
Designed the List → Detail two-panel flow. Specified card anatomy: thumbnail, story-first title, rationale, one-click apply. Added "Review changes" section for transparency. Created interaction specs for hover, dropdown, and back navigation.
Phase 7:
Ran usability sessions validating story-first titles. A/B tested model versions tracking Kept rate. Iterated on "Show details" for power users. Shipped to 10% Fastfood — poor quality dropped 20pp, satisfaction hit 64%.
Initial direction, design and concepts
Three directions explored before converging: auto-apply magic, inline guidance, and story-first previews. User testing killed option one fast.
The MVP
Excel users could insert a chart. Almost none of them could make it good: wrong chart type, unreadable axes, titles that named the chart instead of the finding. About half got deleted in the same session they were created.
I led UX on Copilot Chart Design Recommendations, joining a Redmond-owned engineering track once my other Excel charting work was near done. My job was defining "good" before a model could learn it: a six-dimension rubric, built and enforced with the data science team, that decided what counted as a correct chart recommendation.
The most honest outcome isn't a kept-rate number. The feature reached a 10% fastfood ring, about 5,000 users, before a September 2025 re-org halted the track. It never reached a public rollout, and I don't have production telemetry beyond what I can recall from that ring. What I do have is sharper: our own research caught the shipped system violating my own rubric in the most visible way possible. A user built a pie chart of twelve months of data, and two of our four recommendations were also pie charts. I wrote the rule that bans that. The system didn't enforce it. Finding that, and owning it, is the decision that actually matters here.
43% chart requests to Copilot classified "Generic" — the problem this feature exists to solve
2-4 sec model latency, designed around rather than hidden
9/10 Round 2 participants preferred the recommendation to their own chart (directional, qualitative — not a rate)
2 of 4 recommendations were pie charts for a 12-month dataset — the finding that reset how I read this project's own success
45% of Excel charts got deleted in the same session they were created. That figure moves day to day (45-58% pre-release), so I treat it as "roughly half," not a precise number. Users could insert. They couldn't tell if what they'd inserted was right, and had no path to better except trial and error. "There is no guidance. I have to try five different charts and hope one works," one usability participant put it.
Part of the reason: 43% of chart creation requests to Copilot were generic: "make a chart," with nothing for the model to work with. When people can't say what they want, a model produces something plausible and wrong.
A June 2025 competitive benchmark ran the same visualization prompt across tools: ChatGPT scored 85, Gemini 72, Gemini Sheets 56, BizChat 48. I need to correct something I got wrong in earlier write-ups of this project: that 48 belongs to BizChat, Microsoft's chat surface, not to Copilot in Excel. The real point survives the correction: every external tool in that comparison beat the Microsoft surface being tested, including the weakest of them. That's the case for investing in output quality, not more entry points.
The job to be done was never in question: "Help me create a visualization that tells my story — without needing to be a data viz expert." What people asked for, unprompted, when given a magic wand: automatic chart creation matched to their data, automatic beautification, and natural-language customization. Design Recommendations had to answer all three without pretending to be an expert for them.
This is the one Excel charting project in my set that I joined rather than started. The engineering spine (the shared context pipeline, the on-grid code generation) was owned out of Redmond. I came onto it once my other charting work (Chart Insights) was near done. What I owned inside it was the interaction, the visual system, and the thing the model was missing: a definition of chart quality specific enough that a human annotator and a language model could apply it the same way, every time.
I brought one thing with me from the other project: the on-chart control I'd designed to solve a latency problem on Chart Insights got reused here as the trigger for these recommendations too. One control pattern, two features, one pipeline behind both.
Most AI product design treats the model as a fixed object and designs an interface around it. Here, interface quality depended almost entirely on model quality, and model quality depended on design judgment nobody had written down yet. Writing it down, as a six-dimension rubric, was the actual job.
The rubric had six dimensions: chart type fitness, story-first title, rationale quality, recommendation diversity, data-ink ratio, and whether the data actually met the chosen chart type's conditions. Each came with pass/fail pairs, not abstractions. Chart type fitness: revenue across twelve months routes to a line chart; twelve months of revenue as a pie destroys the temporal sequence. Story-first title: "Q3 Revenue Up 18% vs Q2" passes. "Stacked Bar Chart of Revenue" fails.
Data conditions had their own preconditions per chart type. A line chart needed sequential dates and at least three points. A scatter plot needed two continuous variables and at least ten points. A histogram needed one continuous variable and at least twenty points.
The training set behind all of it ran roughly 60% positive examples, 25% negative, and 15% edge cases, weighted deliberately toward teaching the model what good looked like, not just what to avoid.
Three directions. One got killed fast.
Option 1: AI Takeover. Auto-apply the best recommendation immediately. Users hated it. "Let me see my chart first." That was consistent across every test session. No one wanted the AI to make the call for them. There was a sharper reason underneath the reaction, too: a spreadsheet is a financial and legal record, so an algorithm silently rewriting one is a different kind of decision than an algorithm suggesting a slide edit.
Option 2: Inline tooltips. Minimal guidance surfaced on hover. Too subtle. Didn't break users out of their current habit of guessing.
Option 3: Story-first previews with explicit commit. Show 4 ranked recommendations. Each with a thumbnail, a story-first title ("Quarterly Trends" not "Line Chart"), and a rationale. User picks one. User applies it. User controls the outcome.
Option 3 won because it respected agency while removing effort. Users aren't experts, but they want to feel like they made a decision, not that the machine made it for them.
I debated this with the PM. The instinct was "design-first," lead with the visual. My view was that sequencing trust matters. Users needed to see options before they committed. A preview-first, explicit-commit model built more trust than magic.
Real constraints shaped the rest. LLM latency was 2-4 seconds. Preview fidelity was around 85%, not pixel-perfect. Everything had to live inside a single Copilot pane. These weren't problems to solve. They were facts to design around: a loading state that feels worth the wait, a disclaimer that sets expectations on preview fidelity, four recommendations rather than five (which overwhelmed) or three (which felt too narrow), a refresh button so users could explore more without feeling stuck.
I pushed back on one thing hard: auto-apply. Engineering could have shipped it. It would have been technically impressive. But it would have killed trust. The preview-first model was non-negotiable.
I don't have a clean number to end on, and I'd rather say that than dress one up.
Two moderated research rounds ran in August 2025: five participants each, ten total, qualitative. Round 1 tested flow. Round 2 tested whether the recommendations themselves were any good. Nine of ten participants said they preferred our recommendation to their own chart: encouraging, and not a kept rate. The researcher's own caveat is the honest frame: "Most testers claimed to prefer the applied recommendations over their original charts, but ultimately it's hard to say how true this is given the contrived scenarios and testing environment."
Story-first titles worked. Participants named them, unprompted, as a reason they picked a recommendation. Colour worked mostly as a veto: recommendations got ruled out for low contrast or "doesn't look right on the Excel background," and not one participant across either round chose a dark-mode recommendation willingly.
The finding that mattered came from Round 2. A participant built a pie chart of twelve months of support tickets: twelve slices. Two of our four recommendations were also pie charts. My own rubric bans pie charts above six segments and routes twelve sequential months to a line or column chart. The system did neither. It weighted the user's original chart type heavily enough to override the rule I'd written, then reinforced a choice that was already wrong. That's a failure with my name on it: I wrote the rule and never verified it held at runtime, only in the training set.
On the production side, the FY26 target for Kept/Tried was above 50%. The closest real telemetry in the record - a broader Copilot-in-Excel charting metric from November 2024, not this feature in isolation, and predating this pilot - sat at 24-25.6%. This feature was the attempt to close that gap.
I don't have a telemetry export for what the feature itself did once it shipped - only recollection from the ring it actually ran on. It went to a 2% dogfood ring, about 1,000 users, in early June 2025. Then a 10% fastfood ring, about 5,000 users, that August. It was still there when a September 2025 re-org halted the track. From memory, not measurement: Kept/Tried sat around 65% on that dogfood-biased sample, and the poor-quality rate dropped roughly 20 points, from about 40% to about 20%. Both numbers skew optimistic. A dogfood ring is Microsoft employees, not the broader population the 24-25.6% figure above describes, so I'm not presenting them as apples-to-apples with that number, or as a result I'd stand behind the way I would a measured one. It's what I remember, stated as that. I'd rather describe it accurately than improve it.
Test the runtime against the rubric, not just the training set. I validated examples going into the model and reviewed outputs weekly against the rubric, but I never wrote an automated check that ran the hard constraints against live output. Two pie charts for twelve months of data would have been caught by three lines of code. Instead, a user found it.
Don't assume people read the explanation. I built rationale quality into the rubric as a first-class dimension on the belief it was teaching people something. In the research, most participants skipped it entirely.
Key Design Decisions & Trade-offs
The four biggest calls I made before this reached the dogfood ring — and the reasoning, constraints, and user evidence that shaped each one.
Choice: Generate NATIVE Excel charts, not PNG images
Why: Editable, data-bound, refreshable. Competitors' AI-generated images look good but can't be tweaked.
Choice: Show 1-4 recommendations, prioritized by confidence
Why: Balance guidance with choice. 1 felt prescriptive, 5+ overwhelmed. 4 was sweet spot.
Choice: Thumbnail preview in pane, NOT live chart manipulation on hover
Why: Live preview felt overwhelming. Thumbnails gave control without distraction.
Choice: Always show WHY, not just WHAT to change
Why: Builds user understanding over time. Trust through transparency.
Impact & Results
From dogfood to a 10% fastfood ring, about 5,000 users — before a September 2025 re-org halted the track short of public rollout.
Kept/Tried Rate
Users who apply recommendation keep the chart
Error-Free Load Rate
Down from ~40% pre-pilot — 10% fastfood ring (~5,000 users), recollection, not measured telemetry
Copilot Tried/Enabled Lift
Uplift in users who try Copilot after seeing recommendations
Chart Retention Improvement
Reduction in same-session chart deletions
Feedback from Users
Key finding: Users didn't just apply recommendations — they LEARNED from them. Over time, they started making better initial choices.
"This is like having a data viz expert sitting next to me."
Power User, Internal Preview
"Finally! I don't have to guess if my chart is good."
Intermediate User
"I learned more about charting from these suggestions than from any tutorial."
Novice User, Usability Study
"This is like having a data analyst whispering in my ear when I make a chart."
Financial Analyst, Early Adopter
Internal Testing & Strategic Impact
Pre-launch benchmarks that validated both the technical approach and design decisions before this reached the dogfood ring.
✅ LLM-generated chart code worked every time in controlled tests
✅ Measured against manual chart optimization workflow
✅ All common chart types supported
✅ Column, bar, line, scatter, pie, combo — full MVP coverage
Key Learnings: Designing for AI Collaboration
What I'd do the same, what I'd change, and what this project taught me about designing with — not around — AI.
The Bigger Lesson for AI Product Design