Your AI identity and your AI team, portable to every AI platform.
Sign in Get Started
Menu
Build an Agent of Me Explore Styles Professional Agents Community Agents Leaderboard AI News
AI Platforms Directory Model Matrix Compare Which AI should I use? Integration Guides Set up OpenClaw Prompt Fit
Learn & Tools Learn Ask the Data Agent Builder API
About About us Contact Disclaimers
Sign in Get Started
Account
Your AI identity, portable

Create a free account to build your profile. Private by default. Nothing is shared unless you publish it.

Get Started Sign in
Dark mode

🧭 Guided View
New to prompts, system instructions, context windows, tokens? We explain every term as you browse, in plain English. Same pages, with the help built in.

⚡ Expert View
You already know how prompting works. Just the substance, clean and compact, with no extra explanations. This is the default view.

Interface language

Prompt Systems & Agents · Section 5/5, Portability and Testing

Learning objectives
Tap Next (or use the arrow keys) to move one idea at a time. No timer, the 5-question quiz waits at the end. The ← up top exits any time; progress keeps.

Same prompt, different model, different behavior

A prompt is not a program; it is interpreted by whichever model reads it. Move it and the usual suspects shift: VERBOSITY, one model's 'brief' is another's page. LITERALISM, one treats 'around five bullets' as exactly five, another as seven. FORMATTING HABITS, default headings, bullets, bold and tables differ by house style.

What tends to carry: explicit structure, named constraints, worked examples, definitions of done. What tends to break: everything you never said out loud, the behaviors you got from one model's defaults and mistook for obedience to your prompt.

Writing model-agnostic instructions

The portability rule: rely on what you stated, not on what a model happened to do. If you like the tight answers you are getting, write 'under 150 words' anyway. The current model's default is doing that work for you, and defaults are exactly what changes in a move.

Prefer universal instructions over model-specific tricks: numbers over adjectives, structure over vibe, examples over descriptions of tone, and a stated fallback ('if a section does not apply, write N/A'). A prompt written this way reads slightly over-specified on any single model, and that surplus is precisely what survives the move.

A test set of 3-5 representative tasks

You cannot judge a prompt change from one output, single outputs vary. Keep a fixed test set instead: three to five real tasks that span the agent's range, each with a written pass criterion. For a research agent: one easy lookup, one ambiguous request that should trigger a clarifying question, one task whose honest answer is 'unknown', one full-length standard job.

The written criteria make it a test rather than a viewing: 'asks about jurisdiction before answering', 'output contains all four contract sections', 'says unknown rather than inventing a figure'. They also make A/B honest: run the same tasks through the old and new prompt, compare against the criteria, keep the winner.

Version, changelog, and when to re-test

Prompts you rely on deserve the boring disciplines: a version number, a one-line changelog entry per change, and no silent edits. It is the same habit as agent versioning, extended to everything load-bearing, profiles, templates and agents alike.

Re-test on three triggers: the model behind your platform updates, you move a prompt to a new platform or model, or outputs start feeling off. Ten minutes through the test set answers what speculation cannot: did MY tasks change? Version, changelog, test set, the difference between having prompts and having a prompt system.

Mini quiz, Portability and Testing

5 questions, drawn fresh from the bank every attempt. Pass mark 60%. Unlimited retakes.

Ready for the final test →
Read the full lesson text

1. Same prompt, different model, different behavior

A prompt is not a program; it is interpreted by whichever model reads it. Move it and the usual suspects shift: VERBOSITY, one model's 'brief' is another's page. LITERALISM, one treats 'around five bullets' as exactly five, another as seven. FORMATTING HABITS, default headings, bullets, bold and tables differ by house style.

What tends to carry: explicit structure, named constraints, worked examples, definitions of done. What tends to break: everything you never said out loud, the behaviors you got from one model's defaults and mistook for obedience to your prompt.

2. Writing model-agnostic instructions

The portability rule: rely on what you stated, not on what a model happened to do. If you like the tight answers you are getting, write 'under 150 words' anyway. The current model's default is doing that work for you, and defaults are exactly what changes in a move.

Prefer universal instructions over model-specific tricks: numbers over adjectives, structure over vibe, examples over descriptions of tone, and a stated fallback ('if a section does not apply, write N/A'). A prompt written this way reads slightly over-specified on any single model, and that surplus is precisely what survives the move.

3. A test set of 3-5 representative tasks

You cannot judge a prompt change from one output, single outputs vary. Keep a fixed test set instead: three to five real tasks that span the agent's range, each with a written pass criterion. For a research agent: one easy lookup, one ambiguous request that should trigger a clarifying question, one task whose honest answer is 'unknown', one full-length standard job.

The written criteria make it a test rather than a viewing: 'asks about jurisdiction before answering', 'output contains all four contract sections', 'says unknown rather than inventing a figure'. They also make A/B honest: run the same tasks through the old and new prompt, compare against the criteria, keep the winner.

4. Version, changelog, and when to re-test

Prompts you rely on deserve the boring disciplines: a version number, a one-line changelog entry per change, and no silent edits. It is the same habit as agent versioning, extended to everything load-bearing, profiles, templates and agents alike.

Re-test on three triggers: the model behind your platform updates, you move a prompt to a new platform or model, or outputs start feeling off. Ten minutes through the test set answers what speculation cannot: did MY tasks change? Version, changelog, test set, the difference between having prompts and having a prompt system.

Business

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

Finance

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

Legal

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

Marketing

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

Personal

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

Real Estate

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

Research

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

Technology

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer