> Sample report
What you get back.
Everything on this page is invented. “Example Parcels API” does not exist, and no number below comes from a real client, company or test. It only shows the layout of the reports.
01 Risk report (example)
Every operation in the spec gets a class and a recommended safety label, so the server never treats a risky call as harmless.
| Operation | Class | Suggested label | Decision |
|---|---|---|---|
| GET /parcels | Read-only | read-only | Include |
| GET /parcels/{id} | Read-only | read-only | Include |
| GET /rates | Read-only | read-only | Include |
| GET /parcels/{id}/label | Action disguised as a GET (creates a label and may charge) | costly, not read-only | Hold; needs approval |
| POST /parcels | Create (costly) | write | Hold; not in pilot |
| DELETE /parcels/{id} | Delete | destructive | Exclude |
02 Spec completeness (example)
| Check | Operations |
|---|---|
| Has an operation ID | 5 of 6 |
| Has a description | 4 of 6 |
| Has a success response schema | 6 of 6 |
| Has an example | 2 of 6 |
03 Before and after (example)
30 invented test prompts, 3 repeats of each run. Real reports also list every case whose result changed.
| Measure | Before tuning | After tuning |
|---|---|---|
| Right tool chosen | 17 of 30 | 26 of 30 |
| Right arguments | 14 of 30 | 24 of 30 |
| Correct refusal or question (4 cases) | 1 of 4 | 4 of 4 |
| Calls to a risky tool on a read-only request | 3 | 0 |
Thirty prompts is a small sample, so I report counts rather than headline percentages, and I show how much the results varied between repeats.