blog
The LangWatch Blog
Engineering deep-dives, product updates, and field lessons on evaluating, testing, and observing AI agents in production.
The PR Hound: how we fixed PR review assignment with a daily agent
We leaned into agentic coding and started shipping far more code than our review process could keep up with, so PRs…
Andrew Joia · July 9, 2026
Claude vs Codex: which is the better background agent?
Some of our preferences and stories on Claude and Codex at LangWatch.
Rogerio Chaves · July 8, 2026
Background Agents on Slack: How we built our own Claude Tag before it was cool
A fleet of background agents runs a chunk of our engineering, each one scoped to one job and living in its own Slack…
Rogerio Chaves · July 5, 2026
EU AI Act compliance: are you affected?
The high-risk deadlines just moved to December 2027, but the transparency rules still land on 2 August 2026. What…
Rogerio Chaves · July 5, 2026
Cost per successful task: the metric that will decide your AI stack when model subsidies end
Cost per successful task is run cost divided by success rate. As model subsidies end, it becomes the number that…
Manouk Draisma · July 3, 2026
Introducing: Testing voice agents like you test your chat agents
Test voice agents the way you test text agents - simulated callers, traces, playback, and judge-based evaluation -…
Manouk Draisma · June 2, 2026
The Whole Platform Is Now Open Source: LangWatch May 2026 Update
May 31, 2026 Manouk Draisma
What happens when two engineering teams just... talk
May 19, 2026 Manouk Draisma
LangWatch v3.0 and the April 2026 Product Drop
April 30, 2026 Manouk Draisma
Eat Sleep Append Repeat…
April 20, 2026 Alex Forbes-Reed
Four Refactors and a Funeral: Migrating a Live System to Event Sourcing
April 20, 2026 Alex Forbes-Reed
Internal Product vs Internalised Trauma: Supporting Event Sourced Systems
April 20, 2026 Alex Forbes-Reed
Every way your AI agent can be broken (and how attackers actually do it)
April 15, 2026 Aryan
Why AI Red teaming is broken (and how we fixed it)
April 14, 2026 Rogerio Chaves
How we test Agent Skills with Scenario simulations
March 27, 2026 Sergio Cardenas
Getting to value with LangWatch, faster than ever - how to migrate from Langfuse to LangWatch with Skills.
March 26, 2026 Manouk Draisma
A Note on the LiteLLM Vulnerability
March 25, 2026 Rogerio
Product Managers and leaders are running agent simulations now, and it changing how AI ships
March 25, 2026 Sergio Cardenas
Making your AI Agent reliable: Adding Evaluations to your multi-modal agent with LangWatch Skills
March 24, 2026 Sergio Cardenas
LangWatch Skills: Your coding agent already knows how to test your agent
March 23, 2026 Manouk Draisma
Introducing LangWatch MCP: Test and evaluate AI Agents without leaving your workflow
March 12, 2026 Manouk Draisma
The Agent Development Lifecycle: Why shipping is the easy part
March 6, 2026 Manouk Draisma
The LangWatch February Drop: Cheaper Events, Claude Code, and Multimodal Evals
February 28, 2026 Manouk Draisma
New Pricing: AI growth shouldn’t increase your bill
February 20, 2026 Manouk Draisma
What is LLM monitoring? (Quality, cost, latency, and drift in production)
February 10, 2026 Manouk Draisma
What is Prompt Management? And how to version, control & deploy prompts in productions
February 10, 2026 Manouk Draisma
How OpenClaw / ClawBot works behind the scenes - and why agent observability matters
February 3, 2026 Rogerio Chaves
Instrumenting Your OpenClaw Agent with LangWatch via OpenTelemetry
February 3, 2026 Rogerio Chaves
How to Use Clawdbot + LangWatch to Monitor Your Agents in Production
February 3, 2026 Rogerio Chaves
LLM Evaluations Explained: Experiments, Online Evaluations, Guardrails, and when to use each in 2026
February 2, 2026 Rogerio Chaves
The LangWatch Monthly Drop: January 2026
January 31, 2026 Manouk Draisma
4 best tools for monitoring LLM & agent applications in 2026
January 30, 2026 Bram P
Arize AI alternatives: Top 5 Arize competitors compared (2026)
January 30, 2026 Bram P
Top 10 LLM Observability Tools: Complete Guide for 2026
January 30, 2026 Bram P
Top 5 AI evaluation tools for AI agents & products in production (2026)
January 30, 2026 Manouk Draisma
How to test AI Agents with LangWatch & Mastra / Google ADK and ship them reliably
January 29, 2026 Sergio Cardenas
Top Tools for Evaluating Voice Agents in 2025
December 30, 2025 Bram P
What are the AI Agent Events in 2026: The must-attend conferences for Agentic AI Builders
December 29, 2025 Manouk Draisma
Closing the year Strong: December Product Updates
December 24, 2025 Manouk Draisma
How to do Tracing, Evaluation, and Observability for Google ADK
December 23, 2025 Manouk Draisma
Top 5 AI Prompt Management Tools of 2025
December 23, 2025 Manouk Draisma
Writing Effective AI Evaluations, that hold up in production
December 23, 2025 Manouk Draisma
Why Agentic AI needs a new layer of testing
December 12, 2025 Manouk Draisma
Launch Week Day 5: Better Agents CLI: The reliability layer for the next wave of agent development
November 26, 2025 Rogerio Chaves
Scenario MCP: Automatic Agent Test Generation inside your editor
November 25, 2025 Aryan
Testing Voice Agents with LangWatch Scenario in Real Time
November 24, 2025 Andrew Joia
A Systematic way of Testing of AI Agents
November 20, 2025 Manouk Draisma
Introducing: LangWatch newest Prompt Playground
November 20, 2025 Andrew Garde Joia
How LangWatch helps enterprises test, evaluate, and trust their AI before release
October 27, 2025 Manouk Draisma & FlagSmith
Build vs Buy - Should you build your own LLMOps stack or leverage a purpose-built platform designed for enterprise scale?
October 17, 2025 Manouk Draisma
The 5 Best LLM Evaluation Platforms in 2025: Why LangWatch redefines the category with Agent Testing (with Simulations)
October 17, 2025 Manouk Draisma
Need-based Context Engineering: Let tests tell you what your AI agent actually needs
October 15, 2025 Andrew Joia
The Ultimate RAG Blueprint: Everything you need to know about RAG in 2025/2026
October 6, 2025 Rogerio
From Scenario to Finished: How to Test AI Agents with Domain-Driven TDD
September 26, 2025 Andrew Joia
Building Reliable AI Applications: Why Evals (and Scenarios) Are the backbone of trustworthy AI
September 25, 2025 Manouk Draisma
Are evals dead?
September 7, 2025 Rogerio Chaves
Essential LLM evaluation metrics for AI quality control: From error analysis to binary checks
September 3, 2025 Rogerio Chaves
Trace IDs in AI: LLM Observability and Distributed Tracing
August 22, 2025 Manouk Draisma
The 6 context engineering challenges stopping AI from scaling in production
August 19, 2025 Manouk Draisma
LLMOps is the new DevOps, here’s what every developer must know
August 18, 2025 Manouk Draisma
LLM observability: What is it and why it matters
August 14, 2025 Manouk Draisma
GPT-5 Release: From Benchmarks to production reality
August 8, 2025 Manouk Draisma
LLM-as-a-Judge: Using the Panel of Judges Approach to Approximate Human Preference
August 7, 2025 Rogerio Chaves
Observability Framework Design for LLM Apps - The Complete LangWatch Guide
August 1, 2025 Manouk Draisma
Top 4 Humanloop Alternatives in 2025
July 18, 2025 Manouk Draisma
Why Agent Simulations are the new Unit Tests for AI
July 7, 2025 Tahmid - AI researcher @ LangWatch
Real-time simulation visualization and debug mode
June 27, 2025 Rogerio Chaves
Scripted simulations, evaluations, and guardrails
June 26, 2025 Rogerio Chaves
Customer Story: How Roojoom automates AI Agent Quality Control with LangWatch Scenario
June 25, 2025 Manouk Draisma
Test agents on Mastra, Agno, and 10+ other frameworks
June 25, 2025 Rogerio Chaves
Introducing simulation-based agent testing
June 24, 2025 Rogerio Chaves
Why LangWatch Scenarios represents the future of AI agent testing
June 24, 2025 Rogerio Chaves
Best AI Agent Frameworks in 2025: Comparing LangGraph, DSPy, CrewAI, Agno, and More
June 21, 2025 Rogerio Chaves
Multilingual AI Agent Testing: Using Scenario to Simulate, Break, and Improve LLMs
June 20, 2025 Andrew - Engineer @ LangWatch
LangSmith Alternatives: What to use if you need more security and control
June 18, 2025 Manouk Draisma
Intro to Scenario (Testing AI agents)
June 13, 2025 Tahmid AI researcher @ LangWatch
Simulations from First Principles (How to test your agents)
June 12, 2025 Tahmid AI researcher @ LangWatch
Agent Evaluation: Framework for Testing AI Agents
June 11, 2025 Tahmid AI researcher @ LangWatch
Simulation Based Eval Framework
June 6, 2025 Tahmid - AI research @LangWatch
Introduction: The Real Issue isn’t RL
May 30, 2025 Tahmid AI Researcher @ LangWatch
Simulations to Test My Agent
May 28, 2025 Tahmid, AI Researcher @ LangWatch
New Python SDK Brings Native OpenTelemetry to GenAI Observability
May 15, 2025 Alex Forbes-Reed
April Product Recap: Selene Integration, Eval Wizard Upgrades, Prompt Studio & More
May 5, 2025 Manouk Draisma
LLM Monitoring & Evaluation for Real-World Production Use
May 5, 2025 Manouk Draisma
Systematically Improving RAG Agents
April 24, 2025 Tahmid Tapadar
Introducing the Evaluations Wizard: How to evaluate your LLM: Building an LLM evaluation framework that actually works
April 22, 2025 Rogerio
Function Calling vs. MCP: Why You Need Both - and How LangWatch Makes It Click
April 18, 2025 Manouk Draisma
Why LLM Observability is Now Table Stakes
April 18, 2025 Manouk Draisma
LangWatch vs. LangSmith vs. Braintrust vs. Langfuse: Choosing the Best LLM Evaluation & Monitoring Tool in 2025
April 17, 2025 Manouk Draisma
Introducing Scenario: Use an Agent to Test Your Agent
April 8, 2025 Rogerio Chaves
Tackling LLM Hallucinations with LangWatch: Why Monitoring and Evaluation Matter
April 4, 2025 Manouk Draisma
LLM evaluations at Swis for Dutch government projects by LangWatch
April 3, 2025 Manouk Draisma
Why Your AI Team Needs an AI PM (Quality) Lead
April 2, 2025 Manouk Draisma
LangWatch and adesso join forces: Accelerating Secure LLM Adoption for Enterprises
March 27, 2025 Manouk Draisma
LLMOps Is Still About People: How to Build AI Teams That Don’t Implode
March 25, 2025 Manouk Draisma
Practical LLM Evaluation Framework for AI Development Teams
March 20, 2025 Manouk Draisma
What is Model Context Protocol (MCP)? And how's LangWatch involved?
March 16, 2025 Manouk Draisma
How PHWL.ai uses LLM Observability and Optimization to Improve AI Coaching with LangWatch
March 14, 2025 Manouk Draisma
LangWatch.ai - Announcing - €1M funding round to bring the power of Evaluations and Auto-Optimizations to AI teams.
February 25, 2025 Manouk Draisma
OpenAI, Anthropic, Deepseek and other LLM Providers keep dropping prices: Should you host your own model?
February 20, 2025 Manouk Draisma
7 Predictions for AI in 2025: A CTO's, Rogerio Chaves Perspective
January 1, 2025 Rogerio
Customer Stories: HolidayHero AI start-up <> LangWatch
December 20, 2024 CEO of HolidayHero - redated by Manouk
LangWatch Optimization Studio - Built for AI Engineers, by AI Engineers
December 10, 2024 Rogerio
The power of MIPROv2 (DSPy) in a Low-Code environment with LangWatch’s Optimization Studio
November 10, 2024 Manouk Draisma
What is Prompt Optimization? An Introduction to DSPy and Optimization Studio
November 7, 2024 Manouk Draisma
Deploying an OpenAI RAG Application to AWS ElasticBeanstalk
July 27, 2024 Zhenya
The complete guide for TDD with LLMs
July 3, 2024 Rogerio - CTO
Data Flywheel: Using your production data to build better LLM products
June 27, 2024 Rogerio - CTO
How Algomo reduced AI hallucinations with LangWatch
June 11, 2024 Manouk Draisma
The AI Team: Integrating User and Domain Expert Feedback to Enhance LLM-Powered Applications
June 10, 2024 Manouk Draisma
Unit Testing Your LLM: The Power of Datasets
June 10, 2024 Rogerio Chaves - CTO
Introducing DSPy Visualizer
June 3, 2024 Rogerio - CTO
New Dutch Startup, LangWatch, brings much-needed quality control to GenAI
May 20, 2024 Manouk Draisma
How to build a RAG application from scratch with the least possible AI Hallucinations
May 14, 2024 Zhenya
LLM Reliability with Retrieval-Augmented Generation
May 13, 2024 Manouk Draisma
Safeguarding Your First LLM-Powered Innovation: Essential Practices for Security
May 13, 2024 Manouk Draisma
What is User Analytics for LLMs, The Difference With Traditional Analytics, And Why is it Important?
May 10, 2024 Manouk Draisma
Unlocking the Potential of Large Language Models: The LLM's Beyond the Hype
May 8, 2024 Manouk Draisma
The 8 Types of LLM Hallucinations
May 6, 2024 Manouk Draisma
5 Things You Must Consider Before Putting Your Chatbot Live in Production
May 1, 2024 Manouk Draisma
Navigating the Complexities of AI-Powered Products
May 1, 2024 Manouk Draisma
Understanding Hallucinations: What are they?
April 29, 2024 Manouk Draisma
Mastering the GenAI Wave: Strategies for Success in AI Adoption
April 18, 2024 Manouk Draisma
Successfully building an AI Startup in the current booming industry
April 18, 2024 Manouk Draisma
How Struck.build improved AI Performance with LangWatch
April 17, 2024 Manouk Draisma
Journey Through Innovation: The LLM Adventure
April 8, 2024 Manouk Draisma
Webinar recap: LLM Evaluations: Best Practices, LLM Eval types & real-world insights
Date coming soon Manouk Draisma