The surface needed one worked example on first load — a task the agent had already run, shown end to end. Every candidate we tried demonstrated the product. A flight search proved the agent could act; it did not show you how to tell whether the answer was any good.
Three sessions watching people read a finished run. Nobody read top to bottom. They jumped to the end, found the claim, then went hunting backwards for where it came from. People were reading it as an argument. It was built as a demo.
We took a task with a checkable answer and a wrong turn in the middle — the agent picks up a bad source, notices, backs out. The run is longer and much less impressive. It is also the only version where you can check the ending yourself.
Whether the wrong turn reads as honesty or as a defect the first time someone hits it. No evidence either way until this reaches Stage.