Test Anthropic Claude models with PromptLens
Connect a workspace-owned Anthropic Claude API key and regression-test supported direct model routes against a pinned prompt baseline.
- Updated
- September 30, 2026
- Reading time
- 8 minutes
PromptLens supports direct Anthropic Claude routes in its model catalog. Use this workflow to test a saved prompt or a model migration on your own case set. A model available at the provider is selectable only when its route is supported in PromptLens.
New accounts require early-access approval. A provider API account and usable key are separate prerequisites; a consumer chat subscription does not configure model access in PromptLens.
Key takeaways
Connect at workspace scope
Enter the key in organization provider settings, not in prompt or dataset content.
Record the complete route
Direct and aggregator routes are different evaluated configurations, even for the same model family.
Separate usage from plan price
Generation and optional judge calls can incur provider charges.
1. Connect a direct provider key
Create a key for the intended evaluation environment in the Anthropic Claude developer console. Confirm that it can access the model you intend to use and that billing and usage limits permit a small run. In PromptLens organization provider settings, connect the Anthropic Claude provider.
Use a key intended for this workspace. Keep credentials out of prompt text, uploaded cases, public reports, and logs.
2. Save the baseline prompt and scoring rule
Start with the ten-case triage dataset and its explicit classification instruction. Use User message input mode and Exact match against the reference, trimming whitespace. Save the version, select a supported direct route, and run an initial evaluation. Inspect case failures before using the version as production.
3. Change one variable and compare
For a prompt change, keep the route fixed. For a model migration, keep prompt content, cases, and evaluators fixed and choose the candidate route in a new saved version. Keep supported generation settings comparable; an unavailable setting is not a reason to silently change the experiment.
Compare against the pinned production version. Inspect outputs and explanations for pass-to-fail cases. If an LLM judge is used, explicitly select its route and keep the judge configuration unchanged.
4. Resolve errors before interpreting quality
An authentication failure, unavailable model, rate limit, or provider timeout is an execution failure. It is not evidence that the model violated the rubric. Correct the connection or configuration and rerun affected cases before drawing a quality conclusion.
Review reported token usage and cost when available. Estimates are estimates; missing measurements are not zero cost or zero latency. Provider routing, token limits, and workload size can change the bill.
What this integration covers
This article covers saved text-prompt evaluation through supported direct routes. It does not promise every provider feature, fine-tuning, image generation, batch processing, or arbitrary custom endpoints. A provider capability alone is not an implemented PromptLens capability.