Define acceptable work before starting the clock

Pick one recurring deliverable and name the person who can judge it. For a sales summary, the standard might be: every opportunity is accounted for, every amount matches the source, missing information stays visible, and recommendations are labeled as proposals. A fast draft that fails those checks is unfinished work.

AI delegation still involves human decisions about task fit and responsibility. A study of Microsoft product managers examined those practices through surveys, telemetry, and interviews; it was not a universal productivity benchmark. Primary research. Your own small pilot needs its own task definition, evidence, and reviewer.

  • Use the same output definition and review standard in both phases.
  • Count all per-task human labor once: preparation, active work, checking, and correction.
  • Separate unattended waiting from human effort. Measure elapsed delivery time separately if it matters.
  • Keep setup costs separate and retain failed attempts. Do not select only the examples that look successful.

The worked answer is 56 minutes

The sample contains eight baseline records and eight pilot records with matching task IDs. Baseline work totals 170 minutes. Pilot active work totals 64 minutes, review takes 40, and rework takes 10. So the completed pilot work takes 114 minutes. The fictional reviewer accepts all eight final outputs; three needed correction.

170 − (64 + 40 + 10) = 56 fewer recorded minutes. Baseline average: 21.25 minutes. Pilot average: 14.25 minutes. Difference per pair: seven minutes. The case does not demonstrate that any model is seven minutes faster; it teaches the arithmetic and the measurement boundary.

  • P001: 20 baseline minutes versus 8 active + 5 review + 0 rework = 13 pilot minutes; difference 7.
  • P002: 22 baseline minutes versus 7 active + 5 review + 3 rework = 15 pilot minutes; difference 7.
  • P003: 18 baseline minutes versus 8 active + 6 review + 0 rework = 14 pilot minutes; difference 4.
  • The full answer CSV includes all eight pairs. Ignoring review and correction would overstate the sample difference by 50 minutes.

Prepare two small CSVs that can be reconciled

Start with the provided templates. Use one unique task_id for each comparison case and match it across the baseline and pilot. Keep task descriptions identical only when they describe the same work. An ID match is not evidence that task difficulty, input quality, or reviewer standards were comparable.

The baseline total_minutes field includes all per-task human effort. In the pilot, human_minutes contains preparation and active execution; review_minutes and rework_minutes are separate. Enter accepted, critical_error, and comparable as human judgments. The calculator never opens or grades the actual deliverable.

  • Use yes, no, or unknown for review fields. Blank review decisions remain unknown.
  • Use nonnegative plain minute values with up to two decimals. Missing time is unknown, never zero.
  • Keep source context and explanations in notes. Extra CSV columns are retained in the audit export.
  • Use at most 10,000 data records and 1 MB per input. Everything is processed locally in the browser; no AI service receives the CSVs.

Read the audit before the headline number

The report preserves every input record. It does not invent a zero-effort baseline for unmatched work or select a preferred version of a repeated ID. Both sides of an incomplete pair stay in the audit, with an explicit explanation of why paired arithmetic is unavailable.

Failed outputs still consumed effort. If their times are complete and the work is comparable, those times stay in the comparison. Failed or unknown quality checks block the monthly projection. The same applies to unresolved or excluded source rows, even when the remaining pairs look promising.

  • Duplicate task ID: all versions are withheld, including identical duplicates.
  • Missing review time: the affected pair cannot be calculated; the blank is not replaced with zero.
  • Different task descriptions or unknown comparability: resolve the work definition before comparison.
  • Rejected output or a critical error: retain its time and review the result; a time difference alone is not a business win.
  • Long visible tables are limited to 100 rows. The downloadable audit retains all records and original values.

Separate released capacity from money saved

The optional monthly scenario requires an assumed number of comparable tasks, hourly labor value, and recurring tool cost. A blank is unknown. The calculator withholds the scenario when the evidence or quality review is incomplete. One-time setup is shown separately.

For the fictional seven-minute difference, 200 assumed monthly tasks would represent 23.33 capacity hours. At an assumed $50/hour, that is $1,166.67 of labor-capacity value, or $1,046.67 after an assumed $120 recurring tool cost. The separate $300 setup assumption is reported once. None of those values establishes cash savings, profit, payback, or a staffing recommendation.

A real business would need to show that capacity was usefully redeployed or that actual spending changed. A small comparison also cannot establish causality or statistical significance. Easier pilot tasks, learning effects, different reviewers, and omitted work can all affect the result.

Use the result to design the next useful experiment

In a copy of the sample, add one minute to P001 review time. Pilot effort should become 115 minutes and the total difference 55. Then mark P002 accepted=no: its time should remain included while monthly projection stops. These changes test whether the measurement method handles reality, not just the clean example.

Export the decision memo and choose one next step with your reviewer: continue measurement, revise a slow or unreliable step, or stop this version. Give the next experiment an owner and a review date. If you ask an approved AI tool to summarize the files, require it to separate observations, assumptions, and unresolved evidence. Keep the result a draft and review it before sharing.

The continuing skill is learning how to define work, judge its output, and improve the workflow. SBIH's newsletter carries those lessons into the next task; a calculator result is the beginning of that practice, not permission to automate a whole department.

Before you get started

Does this calculate a realized financial ROI?

It calculates matched recorded effort and, when evidence is complete, an optional capacity-value scenario using your assumptions. It does not establish realized cash savings, profit, a payback period, causality, or a staffing recommendation. Those require additional business evidence.

Does it use AI to grade the pilot?

No. The tool is deterministic. A human enters acceptance, critical-error, and comparability decisions after inspecting the work. It does not run a model, read the final artifact, or validate that those judgments are correct.

What if I do not have a baseline yet?

Use the templates and printable workbook to define the task and start collecting observations. Missing values stay visible. The tool will not turn guessed or blank baseline time into a meaningful comparison.

Are my CSV files uploaded?

The calculator processes their contents locally in your browser, without sending them to an AI model or connecting to your accounts. Downloaded files may still contain your business data. Remove unnecessary private information before sharing an export with anyone else. Newsletter signup is a separate action with explicit consent.

Why does one unresolved task block the monthly projection?

Extrapolating only the clean subset can hide failures or easier selected work. The tool still shows the timed comparable pairs and preserves every excluded record, but asks you to resolve the evidence before projecting monthly capacity. It does not choose an inclusion rule for you.

Sources and how this guide was made

Product guidance is grounded in the sources below. The tools and fictional teaching materials were created for this guide. We do not present these examples as independent product benchmarks or guaranteed outcomes.

Product names belong to their respective owners. Something Big Is Happening is an independent publication. Check current plans, permissions, and availability in the official documentation.