How PromptLens works
Understand the path from an editable prompt to an evaluated release.
A prompt change is useful when it improves the cases you care about. PromptLens keeps the instructions, test cases, results, and published version connected so your team can make that decision with evidence.
From a draft to a release
- Write a prompt. Configure messages, an input mode, and a model route in your personal draft. Each editor has their own draft; editing yours does not change the version an application retrieves.
- Define what good looks like. Add inputs to the prompt's shared dataset and configure evaluators. A reference output supplies the desired answer for checks that need one; other checks use a schema or judging criteria.
- Run and inspect. The initial baseline evaluates the new prompt and establishes saved version 1. Later, save a candidate version and compare it with production. Read individual outputs as well as the overall score.
- Publish deliberately. A label such as
productionpoints to a saved version. After the initial baseline, saving or evaluating a version does not move that label. Assign it to publish; move it back to roll back.
The initial baseline assigns both staging and production to version 1. Completion means the required work finished; it does not mean every assertion passed. Inspect the results before connecting a live application.
What a baseline means
An initial baseline is the first evaluation that completes a new prompt's setup. A production baseline is the production-side result used to compare a later candidate. That comparison pins the production version when it starts, so another release does not change the version being compared mid-run.
Saved versions preserve prompt configuration. Evaluation runs capture the inputs they used. Editing the shared dataset later changes future work, so check a run's coverage before using it to justify a release.
Connect an application
Application retrieval returns the saved messages and settings selected by a label or version number. Your application assembles its input and makes the model call using its own provider credentials.
Observability records traces from a running application so your team can inspect its executions and child model, tool, or retrieval calls. Sending a trace is a separate integration from retrieving a saved prompt.
Work within an organization
Your selected organization owns shared prompts, datasets, evaluation results, traces, model access, and billing. Check the workspace before changing shared resources. CLI and MCP clients have their own organization context.
Continue with Your first evaluation. For exact resource definitions, use the glossary.