Design a decision-ready experiment for [PRODUCT OR USER JOURNEY]. The decision we need to make is [DECISION]. Our evidence about the user problem is [USER EVIDENCE], and the proposed change is [PROPOSED TREATMENT]. Work within these constraints: [CONSTRAINTS AND DATA]. Produce a practical experiment brief with these sections: 1. Decision and hypothesis: state the causal claim in “If we…, then…because…” form; identify the target user population and the exact control and treatment. 2. Success measurement: name one primary metric, its numerator and denominator, measurement window, and why it reflects the decision. Name 2-4 guardrail metrics, including a user-harm or operational-risk guardrail where relevant. Do not treat a proxy as the business outcome without saying so. 3. Design: unit of randomization, eligibility rules, assignment method, exposure event, sample split, duration, and any exclusion rules. Address repeat users, contamination, seasonality, and concurrent changes. 4. Analysis and decision rule: describe the comparison, minimum effect worth acting on, confidence or uncertainty approach available in our tooling, and the precommitted ship, iterate, or stop conditions. Do not promise statistical significance without sample-size inputs. 5. Instrumentation and launch checklist: events, properties, QA cases, ownership, rollback trigger, and an experiment log entry. If the decision is high risk, recommend a staged rollout or qualitative validation first. Do not invent baseline rates, sample size, user research, or feasibility. Ask up to 3 clarifying questions only if a required input is missing. Before answering, check that the treatment can plausibly affect the primary metric and that the proposed measurement cannot be materially biased by assignment, exposure, or missing-event errors. Flag unresolved threats.
Fill in
| Placeholder | What to enter | Example |
|---|---|---|
| [PRODUCT OR USER JOURNEY] | Describe the product area and the user journey being tested. | The first-run dashboard for new team administrators after they connect a data source |
| [DECISION] | State the product or business decision the experiment should inform. | Decide whether to make the setup checklist the default first screen |
| [USER EVIDENCE] | Paste research findings, support themes, funnel data, or observed behavior that motivates the test. | In 22 usability sessions, 15 administrators could not identify the next setup step; only 38% create a first alert within seven days |
| [PROPOSED TREATMENT] | Describe the specific experience or policy change to test. | Show a three-step checklist above the dashboard until the first alert is created |
| [CONSTRAINTS AND DATA] | Enter available traffic, analytics events, risk constraints, engineering limits, and timing. | About 900 eligible admins weekly; events exist for connection, checklist click, alert creation, and seven-day retention; do not delay access to existing dashboard controls |
How to use
- Paste the prompt with a decision that a team can actually make after the test.
- Include the evidence behind the problem so the model does not turn a preference into a hypothesis.
- Confirm event definitions with the analyst or engineer before launch, especially the exposure event.
- Follow up with: “Turn this into a pre-registration note with exact event names and owner assignments.”
Variations
Pricing test
Use this for an experiment involving plans, prices, or paywall presentation.
Design a pricing experiment for [OFFER] among [ELIGIBLE USERS], using [BASELINE AND CONSTRAINTS]. Compare [CONTROL] with [TREATMENT]. Produce a brief covering hypothesis, randomization unit, exposure event, primary metric definition, revenue and conversion guardrails, refund or support-risk guardrails, duration, analysis approach, and ship rule. Address existing customers, currency, taxes, discounts, and price-display compliance. Do not invent baseline conversion or revenue. Ask up to 3 questions only if required information is missing. Flag legal or billing-system review needs.
Qualitative test
Use this before enough traffic exists for a controlled experiment.
Create a moderated usability-test plan for [PROTOTYPE OR FLOW] with [TARGET PARTICIPANTS] to answer [DECISION]. Use [KNOWN USER EVIDENCE]. Produce a 45-minute session guide: screener criteria, realistic scenario, neutral tasks, moderator probes, observation rubric, success signals, risks, and a synthesis template. Do not write leading questions or reveal the intended solution. Ask up to 3 questions only if required inputs are missing. Check that each task tests behavior rather than asking participants to predict what they would do.
Rollout plan
Use this after a test indicates a change may be worth shipping.
Create a staged rollout plan for [CHANGE] based on [EXPERIMENT RESULTS] and [RISK CONSTRAINTS]. Include eligibility, rollout percentages, monitoring cadence, primary and guardrail thresholds, alert owners, rollback criteria, customer communication needs, and the evidence required to advance each stage. Distinguish observed results from assumptions. Do not invent performance or safety data. Ask up to 3 questions only if an input is missing. Check that a rollback can be executed without losing critical user data.
Tips
- Write the decision first; “increase engagement” is not a decision, while “make the checklist the default screen” is.
- Use one primary metric close enough to the change to move during the test, then add guardrails for harms the primary metric could hide.
- Define exposure as the moment a user can actually receive the treatment, not merely when they are assigned to a group.
- Freeze the ship rule before looking at results, or record explicitly why it changed.
FAQ
Can AI determine the sample size for my test?
It can explain the inputs and calculate from supplied baseline, minimum detectable effect, and error tolerance. Have an analyst validate assumptions when the decision is costly or high risk.
What is a guardrail metric?
It is a measure that prevents a local improvement from masking harm, such as support contacts, errors, refunds, latency, or retention.
Should every product change be A/B tested?
No. Use qualitative research for comprehension problems, staged rollouts for risk, and experiments where a randomized comparison can answer the decision.