What Copilot in PowerPoint still can’t do

Every Copilot rollout eventually hits the same moment. Someone sends a generated deck that looks finished and says something wrong, or says the right thing in a way nobody would present. The reaction is usually either “Copilot doesn’t work” or “people need better prompts.” Both are too simple.

Copilot in PowerPoint is genuinely useful. It is also a generative model with specific, predictable limits. Knowing those limits is what lets you decide which work to hand it, which to keep, and which gaps a better template can close. This post lists them plainly.

1. It can’t know what you’re trying to achieve

Copilot works from the prompt and the material you give it. It can structure that material into a presentation. It cannot know that the real goal is to get a budget approved, that the audience is skeptical, or that the most important fact is buried in an appendix.

What closes the gap: a person who knows the intent. Describe the audience and the decision you need in the prompt, and fix the outline before any slides are generated. Microsoft’s workflow lets you do both.

What doesn’t: a better template. Intent is not a design problem.

2. It can’t guarantee accuracy

Microsoft’s documentation says the output of the create-a-presentation feature “is AI-generated content and should be human-reviewed and edited accordingly.” That is the correct stance. Generated slides can summarize incorrectly, drop a qualifier that changes a number’s meaning, or state something with more confidence than the source supports.

What closes the gap: review against the source — every number, name and date — before a deck leaves the building. Brand Skills can add rules like “include a source line under every statistic,” which make review faster. They don’t replace it.

3. It can’t reliably make judgment calls about emphasis

Deciding what matters most on a slide, and making that thing the most visible, is one of the harder skills in presentation design. Copilot applies patterns. When the pattern fits, emphasis is fine. When the content is unusual — a small number that is the headline, a caveat that changes everything — the emphasis often lands in the wrong place.

What closes the gap, partly: example slides in the template that demonstrate your emphasis patterns — highlighted statistics, key takeaways, a single message per slide. Microsoft recommends including layouts for highlighted statistics and key takeaways for exactly this reason. For high-stakes slides, a person’s judgment is still the fix.

4. It can’t do better than the template it learns from

This is the limit organizations most often misdiagnose as a Copilot problem. Microsoft’s guidance is that Copilot learns primarily from a template’s sample slides, uses placeholder type, position and size to decide where content goes, and “falls back to the Slide Master” when there aren’t enough sample slides.

So a template with hardcoded colors, text boxes instead of placeholders and no finished examples produces generic slides, however good the prompt. With strict brand adherence switched on, Copilot can’t even improvise around the gaps.

What closes the gap: rebuilding the template for how Copilot reads it — theme-linked colors and fonts, real placeholders sized for real content, and a set of example slides covering the slide types and densities your people actually use.

5. It can’t read instructions you put in the wrong place

Microsoft is explicit that Copilot does not treat placeholder text as instructions. “Insert three bullets, max ten words each” in a placeholder is a note for a person, not a rule for the model.

What closes the gap: putting rules where Copilot reads them. Since September 2026, brand managers can write note instructions in template speaker notes and publish Brand Skills — written rules — through Brand Kit.

6. It can’t guarantee it follows every rule every time

Even rules in the right place are influence, not code. Strict brand adherence constrains which layouts Copilot uses; note instructions and Brand Skills guide it. A generative model will still occasionally ignore or misapply an instruction.

What closes the gap: testing. A fixed set of realistic prompts, run before and after every change, tells you which rules are working and how often. It also tells you when Microsoft’s changes have altered behavior — which, in September 2026, happened twice.

7. It can’t hold a consistent standard across people on its own

Two employees asking for “a project update” will get different decks, because they wrote different prompts, started from different files and gave different source material. Copilot doesn’t enforce consistency by itself.

What closes the gap: central controls. An official Brand Kit, a single approved template, strict brand adherence, note instructions and Brand Skills all push output toward one standard regardless of who prompts. This is the strongest argument for investing in the template and rules rather than in prompt training alone.

8. It can’t replace the work that carries the most risk

Board presentations, investor decks, keynotes, major pitches, sensitive internal announcements. These are rarely hard because of formatting. They are hard because the stakes are high, the audience is specific and the argument has to be exactly right. Copilot can help draft them. The decisions that make them work are human.

What closes the gap: keeping experienced people — in-house or outside — on the decks that matter, and letting Copilot take the routine ones.

A useful way to sort the work

Put simply, Copilot’s limits fall into three groups:

Limit Fixed by
Generic look, wrong layouts, off-brand output The system: template, Brand Kit, strict mode, notes, skills
Inconsistent behavior, drift over time Testing and oversight: a fixed prompt set, rerun regularly
Intent, accuracy, emphasis, high stakes People: review, judgment, and designers for the decks that matter

The instinct is to invest in the third group through training and review, and to neglect the first. That is backwards. The first group is where a single investment improves every deck every employee generates.

Key takeaways

  • Copilot can’t know your intent, guarantee accuracy or make reliable judgment calls on emphasis.
  • It can’t outperform the template it learns from — the most common cause of generic output.
  • It ignores instructions in placeholders; use note instructions and Brand Skills instead.
  • Rules influence it but don’t bind it, so test and retest.
  • Fix look and consistency with the system, drift with oversight, and intent and stakes with people.

Seeing generic output? Find out whether the template is the cause. Have your template scored.**

Related reading

Sources

24×7 Design Services