From prompt to output: what decides how good AI work gets
How AI systems like Claude actually work, and the seven things that set the quality of what you get back.
Where you run it decides what the AI can use
~/.claude + the folder’s .claude.claude, not your laptopWhat goes into a prompt depends on the job you're asking for
Output quality is a stack, and you control most of it
A language model does one thing: predict the next piece of text
1 · Text is split into tokens
Roughly ¾ of a word each. Models read, think and bill in tokens.
3 · Repeat, token by token
Every word of the answer is produced this way. The model has no memory between chats and no view of your company except what's included in the text it's given.
2 · It scores every possible next token
…frustrated by ___ (illustrative probabilities)
Each turn, the model reads one big bundle, and your prompt is a small part of it
Illustrative split, mid-way through a typical Claude Code session. Modern windows hold a few hundred thousand tokens or more: several books' worth.
A handful of labs build the frontier models
Anthropic
USOpenAI
USGoogle DeepMind
USxAI
USMeta
USDeepSeek
ChinaAlibaba
ChinaMoonshot AI
ChinaZ.ai
ChinaMistral
FranceCapability
Top labs leapfrog each other every few months. Pick a lab for its platform, not just for this month's lead.
Data & compliance
Where data is processed and kept, and whether it's used for training. In a corporation, this usually decides the lab for you.
Closed vs. open weights
Closed: used through the lab's app or API. Open: you can download the model and run it in-house, but you also run and secure it yourselves.
For writing and professional work, the latest scores
Creative writing: which output people prefer
Arena (LMArena): blind side-by-side votes from real users. Best model per lab · Elo · axis starts at 1380.
Professional deliverables (incl. marketing)
GDPval-AA (Artificial Analysis): real tasks from 44 occupations (docs, decks, spreadsheets), judged head to head. Best model per lab · Elo · axis starts at 1450.
Close races are ties. On writing, Gemini 4 Argon and Claude Opus 5.5 are within the margin of error.
Snapshots, not verdicts. Rankings shift every few weeks. Use them as a shortlist.
Your own test beats any chart. Run 5 of your real tasks through 2–3 models and compare.
Three model tiers: pick one for the job, then set how hard it thinks
Quick, cheap, good enough for simple, high-volume jobs.
Near-frontier quality for most tasks at a fraction of the price. The right default.
Deepest judgement and nuance. Opus tops the writing and pro-work rankings; Fable is a premium tier for very long autonomous work.
The second dial: effort / thinking
The same model can answer straight away or reason first. More effort means better answers to hard problems, but it's slower and costs more. Turn it up for analysis and strategy, down for quick edits.
Rule of thumb
Start on the workhorse. Move up when the output lacks judgement or nuance, and down when you're doing the same simple thing hundreds of times.
Anatomy of a strong prompt
Write the context once, and every prompt gets shorter
Without a workspace
Role, audience, voice, rules, examples and format retyped every time, or forgotten and left to guess.
With CLAUDE.md and reference files
Finance leaders, 500–5,000 staff. Pain: slow month-end close.
Senior B2B strategist. Direct, no hype. Never “revolutionary”.
Social posts <150 words, one concrete number.
3 variants in a table, save to drafts/.
Ask up to 3 questions if the brief has gaps.
Read when relevant; CLAUDE.md points to them.
Your prompts now carry only what's new · ~15 words
Specific words change the output; vague words leave it to chance
Swap vague words for checkable ones
Say what to do, not just what not to do. Naming a banned phrase can plant it.
Phrases that change how it works
Ask me questions before you startSurfaces the context you forgot to give
Think this through before answeringMore reasoning on hard or multi-step problems
Give me 3 options that differ in…Breadth instead of one safe answer
Be critical. What's weak here?Counters its natural agreeableness
Only use the attached documentsKeeps it grounded in your sources
If you're unsure, say soFewer confident-sounding guesses
Cite where each claim came fromMakes fact-checking fast
Here's why this matters: …Better judgement on things you didn't specify
Prompting habits that consistently pay off
Brief it like a smart new hire
Brilliant, but on day one. What would they need to know to do this well?
Explain the why
“Readers skim on mobile” beats “use short paragraphs”, and it carries over to choices you didn't spell out.
Show, don't describe
One or two real examples of “good” say more than a paragraph of adjectives.
Material first, question last
With long documents, paste the material at the top and put your ask at the end. Label each part clearly.
Define the output
Length, structure, number of options, table or prose. Don't make it guess the shape.
One job per message
Big tasks go better as a sequence of steps than one giant request.
Let it interview you
“Ask me questions one at a time until you have enough to write this.” Great for briefs and strategy.
Separate draft from critic
After a draft: “Now review it as a sceptical CFO.” Then: “Fix the top 3 issues.”
Ask it to improve your prompt
“Here's my prompt. What's ambiguous or missing?” Prompt-writing is a task it's good at.
Edit, don't argue
If a reply goes wrong, edit your earlier message and re-run instead of piling corrections on top.
Save what works
A prompt that worked becomes a template, then a project instruction, then a skill.
Spot-check against reality
Numbers, names, quotes and links are where errors hide. Verify those first.
Good work comes from a sequence of prompts, not a single one
Brief
Give the situation, audience and goal. Hold back the writing.
Explore
Get options before committing. It's cheap to discard ideas here.
Draft
Commit to one direction with a clear choice.
Critique
Switch the AI from writer to tough reviewer.
Revise
Make targeted fixes, not a full rewrite.
Package
Turn the result into the formats you'll actually use.
Start a fresh chat when…
you switch topics, it keeps repeating the same mistake, or the thread has got very long.
…and carry the essentials over
“Summarise what we've decided and the constraints, as a brief I can paste into a new chat.”
Not every task needs all six
A quick rewrite is just step 3. A strategy doc deserves all six, possibly several times over.
Everything in the thread stays in play, including the wrong turns
The crossed-out turns are still read on every later turn. Their hype keeps pulling the output back.
When context helps
Decisions, preferences and corrections build up. By turn 10 it knows your audience, tone and constraints without being told again.
When context hurts
Rejected drafts, abandoned ideas and unrelated topics compete for attention. Very long threads get summarised automatically, and detail is lost.
How to keep it clean
Edit & re-run the message that went wrong, instead of correcting after it.
One topic per chat. A new task gets a new thread.
Summarise & restart once a thread has done its job.
The workspace briefs the AI before you've typed a word
In Claude Code, the folder you open is the workspace. In the Claude app, Projects work the same way: instructions plus knowledge files.
Loaded every session
- CLAUDE.md files: yours, the team's and the folder's standing instructions
- Memory: facts it saved from earlier sessions
- Names & one-line descriptions of available skills, agents and connectors
Loaded only when needed
- Files it opens or searches
- The full instructions of a skill that matches the task
- Data fetched through connectors
Tools, skills, connectors, plugins and hooks: what each one adds
Tools
Built-in actions: read and write files, search, browse the web, run commands.
Skills
A folder of instructions, templates and examples for one kind of task. Picked up automatically when the task matches.
Connectors
Via MCP, a standard connection protocol. Lets the AI read from and act in other apps.
Plugins
One install that bundles skills, agents, connectors and commands for a role or team.
Hooks
Automatic checks at fixed moments. They run every time, so the AI can't forget them.
Why many skills don't clutter the context: they load in stages
Claude Code can drive other tools, and chain them into workflows
Three ways in: connectors (MCP), APIs called from a short script, and command-line tools. Each one needs your company's approval and its own account or key.
Campaign asset factory
From one brief: headline and copy variants, on-brand images in 4 ad sizes, all saved and named in a folder and sent to Slack for review.
Monday performance digest
Pulls last week's traffic and ad spend, spots what moved and why, makes charts and posts a one-page summary. Runs on a schedule.
Localise a campaign
Adapts the master campaign for 6 markets (not just translation), then generates a native-sounding voiceover for each video cut.
Competitor watch
Visits competitor pricing and feature pages weekly, compares them with last week's snapshot and flags any change worth reacting to.
Personalised outreach
Takes a CRM segment, writes tailored email variants per persona and loads them back as drafts. A human approves before sending.
Design hand-off
Reads a Figma frame, writes copy that fits the actual space, checks it against the brand guide and fills it back into the layout.
An agent is a model working in a loop, with tools and a goal
You ask, it answers once. You do the legwork between turns.
You set a goal. It does the legwork, step by step, then reports back.
A real run in Claude Code
Goal: “Compare our pricing page with our 3 main competitors and draft positioning ideas.”
CLAUDE.md and brand/positioning.mddrafts/pricing-comparison.md with a tableSubagents: helpers with their own fresh context
Clean context
The heavy reading happens in the helper's window. Only the conclusion comes back, so the main thread stays sharp.
Parallel work
Several helpers can run at once: research, analysis and checking happen side by side.
Specialists
Each subagent can have its own instructions, tools and even a different model, like a cheap fast one for bulk reading.
You stay accountable: verify, steer and protect
Check where errors actually happen
- Numbers, dates, names, quotes
- Links and sources: open them
- Claims about your own company or product
- Anything it couldn't actually have seen
AI can be fluent and wrong at the same time. Polished text isn't proof.
Control what agents may do
- Ask first: approve each edit or action
- Plan mode: it researches and proposes, you approve before it acts
- Auto-accept: only for low-risk, reversible work
- Interrupt and redirect at any time
Handle company data properly
- Use the company-approved tool and plan (enterprise plans don't train on your data)
- No customer personal data or secrets in personal accounts
- Connectors inherit your access, so be deliberate
- Label AI-assisted work where policy requires it
When output disappoints, find the layer that failed
Don't just fix the output. Fix whatever produced it.
Helps once. Same problem next time.
Helps every time after, and for your whole team.
Eight habits for better AI output, starting Monday
Match the model to the stakes
Workhorse by default; flagship for judgement calls.
Write the brief once
Standing instructions in a project or CLAUDE.md.
One topic, one thread
Edit instead of arguing; restart with a summary.
Context and the why
Audience, goal, example and format beat clever wording.
Brief, explore, draft, critique
Several focused steps beat one mega-prompt.
Capture what works
Good prompts become templates, then skills.
Delegate goals, not keystrokes
Give an agent a clear “done” and checkpoints.
Trust, then verify
Check the specifics. You own the output.
Sources & further reading
Benchmarks (data as of 2 Oct 2026)
- Arena: Creative Writing leaderboard · arena.ai/leaderboard/text/creative-writing
- Artificial Analysis: GDPval-AA · artificialanalysis.ai/evaluations/gdpval-aa
Prompting
- Anthropic: Prompt engineering overview · docs.claude.com
Claude Code
- Overview & getting started · docs.claude.com/…/claude-code/overview
- Memory & CLAUDE.md · …/claude-code/memory
- Skills · …/claude-code/skills
- Subagents · …/claude-code/sub-agents
- Plugins · …/claude-code/plugins
- MCP connectors · …/claude-code/mcp
- Hooks · …/claude-code/hooks