Guide
How to test a ChatGPT app with record and replay
Record local preview sessions, export reproducible examples and replay interactions. Separate simulated UI checks from real tool and release tests.
Test a complete task, not just a good-looking card
A catalog can look correct while a selected item sends the wrong identifier. A form can render cleanly while its next action loses a field. Test the path from the user’s request to the component interaction and inspect the resolved arguments at each step.
DraftYourApp combines local action simulation with session recording, JSON import/export and playback. Use those features to keep an example that a teammate can reproduce while working on the same interface.
Choose a small set of representative scenarios
- Happy path: a request returns one useful result and the primary action uses the expected ID.
- Empty result: the interface explains that nothing matched and provides a sensible next step.
- Missing input: required information is requested instead of silently invented.
- Follow-up: a filter, form or selected card changes the intended arguments.
Write down the expected visible result and inputs before recording. For an order-status widget, use a synthetic order ID and an expected status. Avoid putting production credentials or private customer records into shareable fixtures.
Record and share a local preview session
- Open the conversation preview and start recording before sending the first message.
- Exercise the relevant components and inspect action feedback and resolved arguments.
- Stop recording and select the saved session. Export its JSON file if you need a reproducible example outside the browser.
- Import a session into the target project and replay it; use the playback controls to inspect the interaction sequence.
Saved sessions are held in the browser’s local storage and are not a shared cloud test repository. Export examples you need to keep or transfer. Treat session files as project data and review them before sharing.
Understand what replay actually proves
Playback replays recorded events without calling external services. Local preview actions simulate tool calls, links and messages rather than executing them. This is useful for investigating a UI sequence, but it does not establish live API availability, a real ChatGPT conversation or model behavior.
Preview can still load visual resources such as images. “No external action execution” should not be read as an entirely offline browser. Run explicit remote tool tests separately when you need evidence about the backend.
Add release checks for supported read-only tools
Per-release regression checks can compare expected structured output with real tool responses for eligible no-auth, read-only, non-destructive GET tools. The check is tied to the immutable release and reports mismatches and tool or transport errors.
An approved sandbox target is a separate explicit choice. It must match the selected tool contracts, and you remain responsible for providing an isolated, authorized endpoint. A matching contract alone does not prove that the sandbox runs the stored bundle.
Keep three kinds of evidence
- Local replay: explains a reproducible interface and interaction sequence.
- Explicit tool or release check: records what the tested endpoint returned.
- ChatGPT host test: checks the actual conversation and widget behavior in its intended environment.
Repeat the relevant checks when you change tool parameters, response schemas or release artifacts. Keep the expected outcome with the example so someone reviewing the next version knows what changed.