LangWatch pricing: free to start, scale by usage
Pricing
Start free. Scale when your agents do.
Start free, add seats at $34 / core-seat / month, and pay only for the events you send. Self-host whenever you want. No credit card to begin.
LangWatch CloudSelf-Managed
Developer
$0 free forever
Everything you need to start building and testing agents.
No credit card required
- 50k events / month
- 14-day data access
- 2 users
- 3 scenarios, 3 simulations, 3 custom evals
- Community support (GitHub & Discord)
Most popular
Growth
$34/ core-seat / month
For teams shipping agents to production.
- Everything in Developer, plus:
- 200k events included, then $6 / 100k
- 30-day retention included (extend at $4 / GB)
- Unlimited lite-users
- Unlimited simulations, evals, prompts
- Private Slack / Teams support
- Volume discounts above 20 users
Enterprise
Custom
For regulated teams that need control and assurance.
- Hybrid, self-hosted or on-prem
- Custom data retention
- Custom SSO / RBAC
- Audit logs & SLAs
- ISO 27001 reports, InfoSec & legal review
- Custom Terms, DPA
- Forward Deployed Engineer
- Billing via AWS / Google Marketplace
Only pay for what you use.
Events at $6 per 100k, on top of $34 / core-seat / month. Storage is billed only when you keep data beyond the included 30-day retention at $4 per GB.
Compare every plan.
| Developer | Growth | Enterprise | |
|---|---|---|---|
| Agent Simulations | Available in all packages | ||
| Simulated users (LLM-powered user simulator) | An AI plays the user, generating realistic messages from your scenario description. | ||
| Multi-turn conversation testing | Test full back-and-forth dialogues, not just single prompts. | ||
| Judge agent, evaluate & verdict at any turn | A judge scores the conversation against criteria and can decide pass/fail at any turn. | ||
| Configurable success criteria | Define what good means in natural language per scenario. | ||
| Scripted to auto-pilot simulations | From fully scripted turns to fully automated runs, choose your level of control. | ||
| Tool-call verification across long dialogues | Assert the right tools were called at the right moments throughout a conversation. | ||
| Framework-agnostic adapters | Works with LangGraph, CrewAI, Pydantic AI and any framework via a one-method adapter. | ||
| Run locally or in CI/CD | Runs in pytest / vitest and in your CI pipeline. | ||
| Simulation visualizer (visual debugging) | Replay and inspect each simulated run step by step. | ||
| Pause, evaluate & annotate mid-conversation | Stop a run at any turn to inspect, score, or annotate. | ||
| Open-source Scenario SDK (Python + TypeScript) | The Scenario testing framework is open source. | ||
| Voice agent testing | End-to-end voice simulations with ElevenLabs, OpenAI Realtime, Twilio, Pipecat, Gemini Live. | ||
| Voice: latency metrics & noise/interruption injection | TTFB, p50/p95 latency, plus background-noise and interruption injection. | ||
| Adversarial / red-teaming | Crescendo escalation, refusal detection and backtracking to surface vulnerabilities. | ||
| Evaluations | |||
| Offline experiments via SDK | Run batch evals over datasets from code. | ||
| Offline experiments via UI | Run and compare experiments with a no-code wizard. | ||
| CI/CD integration | Gate merges on eval results in your pipeline. | ||
| Multi-modal evaluations | Evaluate text, images and more. | ||
| Online evaluations, Monitors | Continuously evaluate production traffic and alert on drops. | ||
| Evaluation by thread | Score whole conversations/threads, not just single messages. | ||
| Guardrails (code integration) | Run evals inline as guardrails in your app. | ||
| Built-in evals | RAGAS, hallucination, toxicity, PII, LLM-as-a-judge and more, out of the box. | ||
| Create reusable evaluators org-wide | Define an evaluator once and share it across projects. | ||
| Custom scoring | Attach your own scores to traces and spans. | ||
| Build custom evals via workflows | Compose evaluators visually in the workflow builder. | ||
| Annotations / annotation inbox | Human-in-the-loop review queue for labeling and feedback. | ||
| Datasets, programmatic access | Create and query evaluation datasets from the API/SDK. | ||
| Generate datasets with AI | Bootstrap datasets automatically with AI. | ||
| Build dataset from traces | Turn real production traces into evaluation datasets. | ||
| Images in datasets | Store and evaluate image inputs in datasets. | ||
| Observability | |||
| Traces and graphs (agents) | Full agent traces with nested spans and the execution graph. | ||
| Session tracking (chats / threads) | Group traces into conversations with a thread id. | ||
| User tracking | Attribute activity to a user id for per-user analytics. | ||
| Topic clustering | Automatically cluster conversations by topic. | ||
| Token and cost tracking | Automatic token and cost accounting per provider, prompt and model. | ||
| Native framework integrations | First-class integrations with popular agent frameworks. | ||
| SDKs (Python, TypeScript) | Official SDKs for Python and TypeScript. | ||
| OpenTelemetry (TypeScript, Go, custom) | OTel-native instrumentation, including Go and custom setups. | ||
| Proxy-based logging (via LiteLLM) | Capture calls through a LiteLLM proxy without code changes. | ||
| Custom via API | Send spans directly via the ingestion API. | ||
| Multi-modal | Capture text, image and audio payloads. | ||
| Additional usage | Events beyond your monthly allowance are billed per 100k. | $6 / 100k | |
| Custom usage pricing | Negotiated volume rates for large deployments. | ||
| Prompt Management | |||
| Prompt version control (code, UI & API) | Versioned prompts editable from code, the UI, or the API. | ||
| Liquid template syntax | Templating with variables and logic via Liquid. | ||
| Prompt data model | Structured prompts with messages, inputs and config. | ||
| Prompt management via GitHub | Sync and review prompts through GitHub. | ||
| Playground | Iterate on prompts side by side across models. | ||
| Prompt experiments / A/B testing | Compare prompt versions on quality, cost and latency. | ||
| Webhooks & Slack | Notify on prompt changes via webhooks and Slack. | ||
| Prompt tags (deployment stages) | Deploy labels (e.g. staging, production) for prompts. | ||
| AI Governance | Enterprise | ||
| Oversight & policy controls | Org-wide control over which assistants, providers and tools each team can use. | ||
| Audit log & CSV export | Every change is logged and exportable to CSV. | ||
| Security | Enterprise | ||
| Custom SSO (Okta, Azure, AWS, Google) | Connect your own identity provider via SAML/OIDC. | ||
| SSO enforcement | Require SSO for all members of your organization. | ||
| Enterprise RBAC (org, project, team) | Granular role-based access control across org, project and team scopes. | ||
| SCIM API for automated user provisioning | Automatically provision and de-provision users from your IdP. | ||
| Audit logs | Full security audit trail of access and actions. | ||
| Data masking | Mask sensitive content in traces and exports. | ||
| Data-retention management | Configure custom retention windows and deletion policies. | ||
| S3 data export (data retention) | Continuously export trace data to your own S3 bucket for long-term retention. | ||
| Compliance | Enterprise | ||
| GDPR / ISO 27001 reports | Compliance reports and documentation on request. | ||
| Custom T&C contracts | Negotiated terms, DPAs and custom contracts. | ||
| InfoSec / legal reviews / PenTest report | Security questionnaires, legal reviews and penetration-test reports. | ||
| Support | |||
| Private Slack / Teams channel | A shared channel with the LangWatch team. | ||
| Onboarding & architectural guidance | Hands-on help designing your setup. | ||
| Dedicated support engineer (deployment & hosting) | A named engineer for deployment and hosting questions. | ||
| Solution architect during evaluation & rollout | Architect support through evaluation and rollout. | ||
| Direct access to the product team | A direct line to product for feedback and requests. | ||
| Billing via AWS / Azure / GCP Marketplace | Consolidate billing through your cloud marketplace. | ||
| Response-time SLO | Guaranteed first-response times. | ||
| Invoice billing | Pay by invoice instead of card. | ||
| Support SLA | Contractual uptime and support service-level agreement. |
Pricing questions, answered.
How does billing work? Billing combines seats and usage. Growth is $34 per core-seat / month; on top you pay for usage, $6 per 100k events beyond the 200k monthly allowance, and $4 per GB only if you retain data beyond the included 30 days. Pay by card, or by invoice and AWS / Azure / GCP marketplace on Enterprise.
How do I track my usage? Your usage, events, storage, and cost by user, prompt and model, is visible live in the dashboard, with anomaly alerts. Enterprise adds an org-wide governance dashboard showing spend by team and top spenders.
What are events? An event is a single ingested span: one LLM call, tool call, or retrieval step within a trace. A typical multi-step agent run produces several events.
How does per-user pricing work? Paid Growth seats are $34 per core-seat / month. Add or remove seats anytime; volume discounts apply above 20 users.
Do I need a credit card to start? No. The Developer plan is free forever, sign up and start sending events with no card required.
Can I self-host? Yes. LangWatch runs fully self-hosted with docker compose on your own ClickHouse, so nothing leaves your environment. Enterprise self-host adds SSO, RBAC, SLAs and support.
Can I change plans or cancel anytime? Yes, upgrade, downgrade or cancel whenever you like.
Try the whole platform, free.
Spin up in minutes on the Developer plan, or talk to us about Growth and Enterprise.