> Sample report
What a pilot report looks like
SAMPLE: invented numbers, invented API (not a real client or company). Every result below is an invented illustration. It is not a measurement, customer evidence, or a claim about actual performance.
Setup
Invented API: “Pebblekite Demo API”, a two-tool server using fake data. A pilot covers up to five tools; this example uses two.
| Tool | Description | Label |
|---|---|---|
| echo | Returns the supplied fake text unchanged. | Read-only |
| add | Adds two supplied numbers without changing stored data. | Read-only |
No write or destructive tools are in this demo. Every pilot tool gets a description and a read-only, write, or destructive label.
Method
30 prompts: 20 for tuning and 10 held back. The held-back set includes four refusal or clarification cases and three ambiguous cases. Each prompt runs three times before and three times after tuning, 180 runs in total. Held-back prompts stay out of tuning. The pilot also includes ChatGPT and Claude connection checks, automated tool calls including an expected failure, and two reviews. This page illustrates the process. It does not document completed testing.
Results (entirely invented)
Tool and argument scores cover cases where a tool should run. Refusal scores cover cases where it should not.
| Set and measure | Before | After |
|---|---|---|
| Tuning: correct tool picked | 45/60 | 58/60 |
| Tuning: correct arguments | 40/60 | 55/60 |
| Held-back: correct tool picked | 12/18 | 16/18 |
| Held-back: correct arguments | 10/18 | 15/18 |
| Held-back: refusals and clarifications handled | 7/12 | 11/12 |
Illustrative checks (also invented): both connections pass; valid automated calls pass; an invalid-number call returns the expected error. Two illustrative reviews cover the tool definitions and the final evidence.
Example failures and fixes
- “Add 8 and 3.” Invented failure: the agent picked echo. Fix: make arithmetic intent clear in the description of add and in the tuning examples.
- “Add 8 to it.” Invented failure: the agent guessed the missing number. Fix: ask for clarification instead of inventing an argument.
Limitations
Real results apply only to the tested model versions and prompts. Testing is not a security certification, and nothing here guarantees future performance, adoption or directory approval.
What you receive in a real pilot
Tool descriptions and labels, the tuning and held-back prompt split, before and after evidence, automated-call checks including the expected failure, and the findings from two reviews.