> Sample report

What a pilot report looks like

SAMPLE: invented numbers, invented API (not a real client or company). Every result below is an invented illustration. It is not a measurement, customer evidence, or a claim about actual performance.

Setup

Invented API: “Pebblekite Demo API”, a two-tool server using fake data. A pilot covers up to five tools; this example uses two.

ToolDescriptionLabel
echoReturns the supplied fake text unchanged.Read-only
addAdds two supplied numbers without changing stored data.Read-only

No write or destructive tools are in this demo. Every pilot tool gets a description and a read-only, write, or destructive label.

Method

30 prompts: 20 for tuning and 10 held back. The held-back set includes four refusal or clarification cases and three ambiguous cases. Each prompt runs three times before and three times after tuning, 180 runs in total. Held-back prompts stay out of tuning. The pilot also includes ChatGPT and Claude connection checks, automated tool calls including an expected failure, and two reviews. This page illustrates the process. It does not document completed testing.

Results (entirely invented)

Tool and argument scores cover cases where a tool should run. Refusal scores cover cases where it should not.

Set and measureBeforeAfter
Tuning: correct tool picked45/6058/60
Tuning: correct arguments40/6055/60
Held-back: correct tool picked12/1816/18
Held-back: correct arguments10/1815/18
Held-back: refusals and clarifications handled7/1211/12

Illustrative checks (also invented): both connections pass; valid automated calls pass; an invalid-number call returns the expected error. Two illustrative reviews cover the tool definitions and the final evidence.

Example failures and fixes

  • “Add 8 and 3.” Invented failure: the agent picked echo. Fix: make arithmetic intent clear in the description of add and in the tuning examples.
  • “Add 8 to it.” Invented failure: the agent guessed the missing number. Fix: ask for clarification instead of inventing an argument.

Limitations

Real results apply only to the tested model versions and prompts. Testing is not a security certification, and nothing here guarantees future performance, adoption or directory approval.

What you receive in a real pilot

Tool descriptions and labels, the tuning and held-back prompt split, before and after evidence, automated-call checks including the expected failure, and the findings from two reviews.