Skip to content
Wednesday, October 7, 2026
iInnovate MagSTARTUPS · INNOVATION · GADGETS · AI
AI

How to Write Better AI Prompts, According to Official Documentation

Prompt engineering is the practice of writing and refining the instructions given to a language model, and the major vendors' official guides agree on the discipline before the tricks: define success criteria, test empirically, then iterate on a draft. Anthropic's documentation covers clarity,…

Mei-Ling Chen · April 22, 2026 · 5 min read
ShareXFacebookLinkedInTelegramEmail
A writer refines a long prompt in a dark editor window, fingertips on keys beneath a warm lamp on a graphite desk.
A writer refines a long prompt in a dark editor window, fingertips on keys beneath a warm lamp on a graphite desk.

Prompt engineering is the practice of writing and refining the instructions given to a language model, and the major vendors' official guides agree on the discipline before the tricks: define success criteria, test empirically, then iterate on a draft. Anthropic's documentation covers clarity, examples, XML structuring, role prompting, thinking, and prompt chaining as its core techniques, all documented publicly.

What do the official guides agree on?

Both Anthropic's and Google's guides treat prompting as an engineering loop rather than incantation. Anthropic's overview page frames the precondition honestly: "This guide assumes that you have: A clear definition of the success criteria for your use case Some ways to empirically test against those criteria A first draft prompt you want to improve If not, spend time establishing that first," per the prompt engineering overview in Claude's documentation. The writing comes third, after measurement exists.

Google's Gemini documentation takes the same empirical stance and spends its pages on structure: system instructions, format specification, and strategies for steering a response before it begins. Neither guide promises magic words, and both implicitly concede the opposite: a prompt that cannot be tested cannot be improved, only reworded.

How should a prompt be structured, per Claude's documentation?

Anthropic's technique list is the backbone, and each entry answers a specific failure mode. Clarity means instructions a competent new hire could follow without guessing. Examples demonstrate the desired output instead of describing it, which removes an entire class of format surprises. XML structuring separates instructions from data with explicit tags, so the model can tell which words are orders and which are material.

The remaining three techniques scale the reasoning. Role prompting establishes who the model should act as, which constrains vocabulary and defaults. Thinking reserves space for the model to work through a problem before committing to an answer. Prompt chaining decomposes a complex job into a sequence of smaller, individually checkable prompts, which converts one unmanageable failure into several findable ones.

Applied selectively they compose into a documented workflow: a role, a task in tags, two worked examples, and a chain of steps covers most real jobs better than a single overloaded paragraph ever does.

What strategies does Google document for Gemini?

Google's prompt design guide for the Gemini API concentrates on controlling the response before generation starts. One documented section addresses output shape directly: "You can give instructions that specify the format of the response. For example, you can ask for the response to be formatted as a table, bulleted list, elevator pitch, keywords, sentence, or paragraph," per the prompt design strategies in Google's documentation. Format instructions move decisions out of the model's discretion and into the prompt where they belong.

The same guide documents system instructions as a separate channel that persists across turns. Its worked example instructs the model to answer comprehensively with detail unless the user requests a concise response, which shows the pattern: a standing rule in the system layer, task-specific content in the prompt layer. Google also documents the completion strategy, in which the prompt supplies the beginning of the desired output and the model continues it, a technique that also helps enforce a format.

The guide's own worked examples, including a sample question about starting a DVD business in 2026 answered by gemini-2.5-flash, are worth reading in full, because they show the vendors testing their own advice in public rather than asserting it.

What is the documented workflow for improving a prompt?

Assembled from both vendors' pages, the loop is short and repeatable:

  1. Define success criteria for the task before writing anything, per Anthropic's overview.
  2. Build a way to test empirically against those criteria, even a small fixed set of cases.
  3. Write a first draft prompt and run it against the tests.
  4. Change one variable at a time: format, role, examples, or structure.
  5. Chain the task apart when a single prompt keeps failing, and re-test each stage.

The discipline in steps four and five is what separates iteration from thrash. A prompt with three simultaneous changes gives no evidence about which change mattered, and the guides' insistence on empirical testing exists precisely to prevent that ambiguity.

What can't prompting fix?

The guides are candid at the margins. Anthropic's overview is framed around learning when prompt engineering is the right solution at all, implying a category of problems it is not: unclear goals, unmeasurable outputs, and tasks the underlying model cannot do regardless of wording. Better instructions remove interference; they do not add capability.

Diagnosing which of the two is failing is the actual skill the documentation teaches. If a clean, chained, well-formatted prompt still fails on a fixed test set, the issue is the model or the task framing, not the adjectives, and no template will paper over it.

There is also a documented middle category: tasks the model can do but only with the right structure. Long documents that fail as one giant prompt often pass when split across a chain; outputs that drift in tone tighten with a system instruction; answers that miss required fields comply when the format is specified up front. These are the cases where the guides genuinely pay off, and they are identifiable only because the testing loop flagged them in the first place.

The compounding effect is organizational, not personal. A team that writes its success criteria and test cases once ends up with a reusable evaluation harness, and every later prompt, model swap, or provider comparison runs against the same yardstick. That artifact outlives any individual prompt, which is the quiet argument both vendors' guides are making.

Sources

  1. Prompt engineering overview — Anthropic
  2. Prompt design strategies — Google

More from our brands

Part of the VUGA Network