# langwatch.ai > AI-optimized mirror of langwatch.ai containing 50 pages totalling 45,694 words of clean markdown content, structured data, and semantic HTML. Original source: https://langwatch.ai. Last updated: 2026-07-20T14:38:12.861Z. Each page is available as HTML (with JSON-LD structured data) and Markdown (text-only, ideal for LLMs and RAG). ## Homepage - [AI Agent Testing and Evaluation | LangWatch](/content/site-root.html): AI agent testing and evaluation that turns unpredictable agents into reliable production systems, with simulations, evals, observability, and governance. (817 words) ## Articles & Blog Posts - [blog/pms-and-ceos-are-running-agent-simulations-now-e2-80-94-and-it-s-changing-how-ai-ships.html](/content/blog/pms-and-ceos-are-running-agent-simulations-now-e2-80-94-and-it-s-changing-how-ai-ships.html) (1,435 words) - [Every Way Your AI Agent Can Be Broken by Attackers](/content/blog/every-way-your-agent-can-be-broken/index.html): A practical AI agent red teaming guide: attacker goals, the Crescendo multi-turn strategy, OWASP LLM Top 10 mapping, and how to test agents before shipping. (2,379 words) - [blog/index.html](/content/blog/index.html) (2,105 words) - [Migrating a Live System to Event Sourcing](/content/blog/four-refactors-and-a-funeral-migrating-a-live-system-to-event-sourcing.html): How LangWatch migrated a live platform to event sourcing on ClickHouse: four refactors, a fold-map-react pipeline, and a zero-downtime dual-write cutover. (4,334 words) - [Why AI Red Teaming Is Broken and How We Fixed It](/content/blog/why-ai-red-teaming-is-broken-and-how-we-fixed-it/index.html): Single-turn AI red teaming misses real attacks. See how multi-turn Crescendo simulations in Scenario break agents that pass every benchmark. (2,147 words) - [EU AI Act Compliance: Are You Affected? 2026 Guide](/content/blog/eu-ai-act-compliance-llm-applications/index.html): Are you affected by the EU AI Act? A decision tree for LLM apps, the amended 2026-2028 timeline, deployer duties, logging, oversight, and a Q3 checklist. (3,652 words) - [Background Agents on Slack: How We Built Our Own Claude Tag](/content/blog/background-agents-before-claude-tag/index.html): Seven background agents run our own engineering, each scoped to one job in its own Slack channel, months before Anthropic shipped the pattern as Claude Tag. (2,204 words) - [LangWatch pricing: free to start, scale by usage](/content/pricing/index.html): Start free with the LangWatch Developer plan. Paid plans from €29/month add unlimited evaluations and simulations, with enterprise security and self-hosting. (1,355 words) - [Migrate From Langfuse to LangWatch With Skills](/content/blog/getting-to-value-with-langwatch-faster-than-ever-how-to-migrate-to-langwatch-with-skills.html): Set up evals, agent simulations, and migrate from Langfuse or LangSmith to LangWatch in one session using Skills and MCP, not a full sprint of work. (1,351 words) - [Ops Tooling for Event-Sourced Systems in Production](/content/blog/internal-product-vs-internalised-trauma-supporting-event-sourced-systems.html): The tooling that keeps an event-sourced system supportable: Grafana dashboards, an ops console for blocked groups, projection replay, and time-travel… (1,582 words) - [Testing Voice Agents Like You Test Chat Agents](/content/blog/introducing-testing-voice-agents-like-you-test-your-chat-agents.html): Test voice agents with Scenario using simulated callers, traces, playback, and judge-based evaluation against real voice across OpenAI, Twilio, and more. (2,084 words) - [How We Test Agent Skills With Scenario Simulations](/content/blog/how-we-test-agent-skills-with-scenario-simulations/index.html): Most AI agent skills ship untested. See how LangWatch Scenario simulates full multi-turn conversations with a user simulator and judge to catch real failures. (1,456 words) - [Cost per Successful Task: Pricing AI Agents After Subsidies](/content/blog/cost-per-successful-task/index.html): Cost per successful task = run cost ÷ success rate. Learn how to calculate it for your AI agents and switch to cheaper models before flat-rate plans end. (1,467 words) - [Event Sourcing Made LangWatch 300x Faster](/content/blog/eat-sleep-append-repeat-e2-80-a6/index.html): Re-architecting LangWatch on event sourcing and ClickHouse hit 6,000 events per second, real-time dashboards, and retroactive data improvements via replay. (1,298 words) - [Changelog | LangWatch](/content/changelog/index.html): Every LangWatch release, improvement, and fix as it ships. Follow the platform evolve across evals, simulations, observability, and prompt management. (656 words) - [AI Governance | LangWatch](/content/gateway/index.html): Govern every model, key, and tool through one AI gateway. Virtual keys with budgets, routing policies, fallback, and a full audit trail. Cloud or self-hosted. (908 words) - [Integrations | LangWatch](/content/integrations/index.html): Every integration LangWatch ships: Python and TypeScript SDKs, OpenTelemetry, OpenAI, Anthropic, AWS Bedrock, Azure, Vertex, LangGraph, CrewAI, DSPy, and more. (482 words) - [A Note on the LiteLLM Vulnerability and LangWatch](/content/blog/a-note-on-the-litellm-vulnerability/index.html): A malicious credential-stealing file appeared in LiteLLM 1.82.8. LangWatch customers were not impacted because all dependencies are pinned with lockfiles. (426 words) - [The AI Agents Guide — LangWatch](/content/guides/ai-agent-guide/index.html) (1,055 words) - [Codex CLI - LangWatch](/content/docs/ai-gateway/cli/codex/index.html): Route the OpenAI Codex CLI through the LangWatch AI Gateway. (787 words) - [Voice agent testing | LangWatch](/content/voice-ai-agent-testing/index.html): At-scale automated testing for voice and chat agents. Open-source simulation, latency and quality checks, and production monitoring for voice AI. (774 words) - [LangWatch Goes Open Source: May 2026 Update](/content/blog/langwatch-monthly-drop-may-2026/index.html): LangWatch is now open source under Apache 2.0. Plus native voice testing in Scenario, the GA Traces UI, and the new AI Governance beta. See what shipped in… (809 words) - [Track your Codex usage | LangWatch](/content/codex-usage/index.html): Full trace history for Codex sessions: every model turn and tool call, tokens by class including cache reads, and theoretical spend on bundled ChatGPT plans. (520 words) - [Prompt management | LangWatch](/content/prompt-management/index.html): Manage prompts as one versioned source of truth. Version, review through GitHub, A/B test in production, and tie every prompt to its traces. Prompts as code. (543 words) - [The Evals Golden Guide — LangWatch](/content/guides/evals-guide/index.html) (21 words) - [LLM red-teaming for AI agents | LangWatch](/content/llm-red-teaming/index.html): Open-source adversarial testing for AI agents. Run 50-turn Crescendo attacks to surface jailbreaks, prompt extraction, and data exfiltration before they ship. (526 words) - [LangWatch vs LangSmith vs LangFuse | Comparison](/content/comparison/index.html): Compare LangWatch, LangSmith, and LangFuse across LLM observability, evaluations, guardrails, and production readiness, feature by feature. (242 words) - [LLM observability | LangWatch](/content/llm-observability/index.html): OpenTelemetry-native LLM observability. Trace every agent step, monitor cost and latency, and debug AI in production with full end-to-end visibility. (549 words) - [Agent simulation testing | LangWatch](/content/agentic-ai-testing/index.html): Test AI agents with scenario simulations before they reach production. Simulate real conversations, catch regressions, and prove agent quality every release. (501 words) - [Self-Hosting Overview - LangWatch](/content/docs/self-hosting/overview/index.html): Deploy LangWatch on your own infrastructure for full data control (485 words) - [LangWatch vs LangSmith vs Braintrust vs Langfuse](/content/blog/langwatch-vs-langsmith-vs-braintrust-vs-langfuse-choosing-the-best-llm-evaluation-monitoring-tool-in-2025.html): Compare LangWatch, LangSmith, Braintrust, and Langfuse to choose the best LLM evaluation and monitoring tool in 2025, with the capabilities that matter most. (702 words) - [LLM evaluation | LangWatch](/content/llm-evaluation/index.html): Evaluate LLM and agent quality with offline and online evals, custom evaluators, and datasets. Catch quality regressions before your users ever do. (421 words) - [Arize alternative | LangWatch](/content/arize-alternative/index.html): Built for LLMs, not retrofitted from classic ML monitoring. The Arize alternative with agent simulation testing, prompt optimization, and a hybrid workflow. (575 words) - [Partner program | LangWatch](/content/partners/index.html): Join the LangWatch partner program. Refer, resell, or integrate, and bring reliable AI to your customers with revenue share up to 45% and sales support. (314 words) - [SDKs | LangWatch](/content/sdks/index.html): LangWatch Python and TypeScript SDKs for tracing, evaluating, and optimizing agents. OpenTelemetry-native and framework-agnostic. Quickstart in five minutes. (206 words) - [LLM Nodes - LangWatch](/content/docs/optimization-studio/llm-nodes/index.html): Use LLM Nodes in Optimization Studio to invoke LLMs from workflows and run controlled evaluations for agent testing. (512 words) - [The Better Agents Manifesto | LangWatch](/content/better-agents-manifesto/index.html): Four values for teams building agents that ship. Systematic quality, real scenarios, business metrics, incremental over premature AGI. (736 words) - [LangWatch Launch Week | Five days of releases](/content/launch-week-nov/index.html): Five days, five releases: Prompt Playground, RBAC, voice agent testing, Scenario MCP, and the Better Agents CLI. Everything from LangWatch Launch Week. (304 words) - [Customer stories | LangWatch](/content/customers/index.html): How AI teams use LangWatch to evaluate, monitor, and ship reliable agents in production. Read the customer stories and the results they reached. (313 words) - [Track your Claude Code usage | LangWatch](/content/claude-code-usage/index.html): Full trace history for Claude Code, Codex, OpenCode, Copilot, Cursor, and Pi. Tokens by class including cache reads, and theoretical spend on bundled plans. (341 words) - [Track your OpenCode usage | LangWatch](/content/opencode-usage/index.html): Full trace history for OpenCode sessions on any provider: model turns, tool calls, tokens by class including cache reads, and theoretical spend per model. (308 words) - [Book a LangWatch demo | Talk to the team](/content/get-a-demo/index.html): See LangWatch on your stack: real evaluations and agent simulations, pricing, security, and procurement. Book a 30-minute session with our team. (353 words) - [Humanloop alternative | LangWatch](/content/humanloop-alternative/index.html): Multi-turn, multi-tool, open source. The Humanloop alternative with full agent simulation, OpenTelemetry-native tracing, and built-in red teaming. (385 words) - [Trust center | LangWatch security & compliance](/content/trust-center/index.html): LangWatch security and compliance: ISO 27001, GDPR, SOC 2, encryption, RBAC, SSO, and self-hosting. Read our trust report and deployment options. (233 words) - [AI agent testing training for your team | LangWatch](/content/ai-agent-testing-training/index.html): A guided curriculum for teams evaluating LLM agents. Six modules from scoring to scenario design and CI gates. Workshop or self-paced. (287 words) - [Thumbs Up/Down - LangWatch](/content/docs/user-events/thumbs-up-down/index.html): Track thumbs up/down user feedback in LangWatch to evaluate LLM quality and guide AI agent testing improvements. (231 words) - [Braintrust alternative | LangWatch](/content/braintrust-alternative/index.html): Stop scoring failures. Start preventing them. LangWatch is the Braintrust alternative with agent simulation, open source, and a workflow for the whole team. (287 words) ## Listings & Categories - [Aryan Sharma | LangWatch Blog](/content/blog/author/aryan-sharma/index.html): 2 posts by Aryan Sharma, Engineer at LangWatch on testing, evaluating, and observing AI agents in production. (80 words) - [Tahmid Tapadar | LangWatch Blog](/content/blog/author/tahmid-tapadar/index.html): 8 posts by Tahmid Tapadar, AI Researcher, PhD at LangWatch on testing, evaluating, and observing AI agents in production. (156 words) ## Resources - [Full Page Index](/index.html): Browse all cached pages with rich metadata - [About This Cache](/content/about.html): Methodology, technical details, and usage guidelines - [XML Sitemap](/sitemap.xml): Machine-readable sitemap for crawler discovery - [Robots.txt](/robots.txt): Crawler directives