# The LangWatch Blog

Engineering deep-dives, product updates, and field lessons on evaluating, testing, and observing AI agents in production.

## The PR Hound: how we fixed PR review assignment with a daily agent

We leaned into agentic coding and started shipping far more code than our review process could keep up with, so PRs…

[Andrew Joia](/content/blog/author/andrew-joia/index.html) · July 9, 2026

## Claude vs Codex: which is the better background agent?

Some of our preferences and stories on Claude and Codex at LangWatch.

[Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html) · July 8, 2026

## Background Agents on Slack: How we built our own Claude Tag before it was cool

A fleet of background agents runs a chunk of our engineering, each one scoped to one job and living in its own Slack…

[Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html) · July 5, 2026

## EU AI Act compliance: are you affected?

The high-risk deadlines just moved to December 2027, but the transparency rules still land on 2 August 2026. What…

[Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html) · July 5, 2026

## Cost per successful task: the metric that will decide your AI stack when model subsidies end

Cost per successful task is run cost divided by success rate. As model subsidies end, it becomes the number that…

[Manouk Draisma](/content/blog/author/manouk-draisma/index.html) · July 3, 2026

## Introducing: Testing voice agents like you test your chat agents

Test voice agents the way you test text agents - simulated callers, traces, playback, and judge-based evaluation -…

[Manouk Draisma](/content/blog/author/manouk-draisma/index.html) · June 2, 2026

### The Whole Platform Is Now Open Source: LangWatch May 2026 Update

May 31, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### What happens when two engineering teams just... talk

May 19, 2026  [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### LangWatch v3.0 and the April 2026 Product Drop

April 30, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Eat Sleep Append Repeat…

April 20, 2026 [Alex Forbes-Reed](/content/blog/author/alex-forbes-reed/index.html)

### Four Refactors and a Funeral: Migrating a Live System to Event Sourcing

April 20, 2026 [Alex Forbes-Reed](/content/blog/author/alex-forbes-reed/index.html)

### Internal Product vs Internalised Trauma: Supporting Event Sourced Systems

April 20, 2026 [Alex Forbes-Reed](/content/blog/author/alex-forbes-reed/index.html)

### Every way your AI agent can be broken (and how attackers actually do it)

April 15, 2026 [Aryan](/content/blog/author/aryan-sharma/index.html)

### Why AI Red teaming is broken (and how we fixed it)

April 14, 2026 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### How we test Agent Skills with Scenario simulations

March 27, 2026 [Sergio Cardenas](/content/blog/author/sergio-cardenas/index.html)

### Getting to value with LangWatch, faster than ever - how to migrate from Langfuse to LangWatch with Skills.

March 26, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### A Note on the LiteLLM Vulnerability

March 25, 2026 [Rogerio](/content/blog/author/rogerio-chaves/index.html)

### Product Managers and leaders are running agent simulations now, and it changing how AI ships

March 25, 2026 [Sergio Cardenas](/content/blog/author/sergio-cardenas/index.html)

### Making your AI Agent reliable: Adding Evaluations to your multi-modal agent with LangWatch Skills

March 24, 2026 [Sergio Cardenas](/content/blog/author/sergio-cardenas/index.html)

### LangWatch Skills: Your coding agent already knows how to test your agent

March 23, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Introducing LangWatch MCP: Test and evaluate AI Agents without leaving your workflow

March 12, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### The Agent Development Lifecycle: Why shipping is the easy part

March 6, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### The LangWatch February Drop: Cheaper Events, Claude Code, and Multimodal Evals

February 28, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### New Pricing: AI growth shouldn’t increase your bill

February 20, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### What is LLM monitoring? (Quality, cost, latency, and drift in production)

February 10, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### What is Prompt Management? And how to version, control & deploy prompts in productions

February 10, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### How OpenClaw / ClawBot works behind the scenes - and why agent observability matters

February 3, 2026 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Instrumenting Your OpenClaw Agent with LangWatch via OpenTelemetry

February 3, 2026 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### How to Use Clawdbot + LangWatch to Monitor Your Agents in Production

February 3, 2026 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### LLM Evaluations Explained: Experiments, Online Evaluations, Guardrails, and when to use each in 2026

February 2, 2026 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### The LangWatch Monthly Drop: January 2026

January 31, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### 4 best tools for monitoring LLM & agent applications in 2026

January 30, 2026 [Bram P](/content/blog/author/bram-p/index.html)

### Arize AI alternatives: Top 5 Arize competitors compared (2026)

January 30, 2026 [Bram P](/content/blog/author/bram-p/index.html)

### Top 10 LLM Observability Tools: Complete Guide for 2026

January 30, 2026 [Bram P](/content/blog/author/bram-p/index.html)

### Top 5 AI evaluation tools for AI agents & products in production (2026)

January 30, 2026 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### How to test AI Agents with LangWatch & Mastra / Google ADK and ship them reliably

January 29, 2026 [Sergio Cardenas](/content/blog/author/sergio-cardenas/index.html)

### Top Tools for Evaluating Voice Agents in 2025

December 30, 2025 [Bram P](/content/blog/author/bram-p/index.html)

### What are the AI Agent Events in 2026: The must-attend conferences for Agentic AI Builders

December 29, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Closing the year Strong: December Product Updates

December 24, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### How to do Tracing, Evaluation, and Observability for Google ADK

December 23, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Top 5 AI Prompt Management Tools of 2025

December 23, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Writing Effective AI Evaluations, that hold up in production

December 23, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Why Agentic AI needs a new layer of testing

December 12, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Launch Week Day 5: Better Agents CLI: The reliability layer for the next wave of agent development

November 26, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Scenario MCP: Automatic Agent Test Generation inside your editor

November 25, 2025 [Aryan](/content/blog/author/aryan-sharma/index.html)

### Testing Voice Agents with LangWatch Scenario in Real Time

November 24, 2025 [Andrew Joia](/content/blog/author/andrew-joia/index.html)

### A Systematic way of Testing of AI Agents

November 20, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Introducing: LangWatch newest Prompt Playground

November 20, 2025 [Andrew Garde Joia](/content/blog/author/andrew-joia/index.html)

### How LangWatch helps enterprises test, evaluate, and trust their AI before release

October 27, 2025 [Manouk Draisma & FlagSmith](/content/blog/author/manouk-draisma/index.html)

### Build vs Buy - Should you build your own LLMOps stack or leverage a purpose-built platform designed for enterprise scale?

October 17, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### The 5 Best LLM Evaluation Platforms in 2025: Why LangWatch redefines the category with Agent Testing (with Simulations)

October 17, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Need-based Context Engineering: Let tests tell you what your AI agent actually needs

October 15, 2025 [Andrew Joia](/content/blog/author/andrew-joia/index.html)

### The Ultimate RAG Blueprint: Everything you need to know about RAG in 2025/2026

October 6, 2025 [Rogerio](/content/blog/author/rogerio-chaves/index.html)

### From Scenario to Finished: How to Test AI Agents with Domain-Driven TDD

September 26, 2025 [Andrew Joia](/content/blog/author/andrew-joia/index.html)

### Building Reliable AI Applications: Why Evals (and Scenarios) Are the backbone of trustworthy AI

September 25, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Are evals dead?

September 7, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Essential LLM evaluation metrics for AI quality control: From error analysis to binary checks

September 3, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Trace IDs in AI: LLM Observability and Distributed Tracing

August 22, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### The 6 context engineering challenges stopping AI from scaling in production

August 19, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### LLMOps is the new DevOps, here’s what every developer must know

August 18, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### LLM observability: What is it and why it matters

August 14, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### GPT-5 Release: From Benchmarks to production reality

August 8, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### LLM-as-a-Judge: Using the Panel of Judges Approach to Approximate Human Preference

August 7, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Observability Framework Design for LLM Apps - The Complete LangWatch Guide

August 1, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Top 4 Humanloop Alternatives in 2025

July 18, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Why Agent Simulations are the new Unit Tests for AI

July 7, 2025 [Tahmid - AI researcher @ LangWatch](/content/blog/author/tahmid-tapadar/index.html)

### Real-time simulation visualization and debug mode

June 27, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Scripted simulations, evaluations, and guardrails

June 26, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Customer Story: How Roojoom automates AI Agent Quality Control with LangWatch Scenario

June 25, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Test agents on Mastra, Agno, and 10+ other frameworks

June 25, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Introducing simulation-based agent testing

June 24, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Why LangWatch Scenarios represents the future of AI agent testing

June 24, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Best AI Agent Frameworks in 2025: Comparing LangGraph, DSPy, CrewAI, Agno, and More

June 21, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Multilingual AI Agent Testing: Using Scenario to Simulate, Break, and Improve LLMs

June 20, 2025 [Andrew - Engineer @ LangWatch](/content/blog/author/andrew-joia/index.html)

### LangSmith Alternatives: What to use if you need more security and control

June 18, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Intro to Scenario (Testing AI agents)

June 13, 2025 [Tahmid AI researcher @ LangWatch](/content/blog/author/tahmid-tapadar/index.html)

### Simulations from First Principles (How to test your agents)

June 12, 2025 [Tahmid AI researcher @ LangWatch](/content/blog/author/tahmid-tapadar/index.html)

### Agent Evaluation: Framework for Testing AI Agents

June 11, 2025 [Tahmid AI researcher @ LangWatch](/content/blog/author/tahmid-tapadar/index.html)

### Simulation Based Eval Framework

June 6, 2025 [Tahmid - AI research @LangWatch](/content/blog/author/tahmid-tapadar/index.html)

### Introduction: The Real Issue isn’t RL

May 30, 2025 [Tahmid AI Researcher @ LangWatch](/content/blog/author/tahmid-tapadar/index.html)

### Simulations to Test My Agent

May 28, 2025 [Tahmid, AI Researcher @ LangWatch](/content/blog/author/tahmid-tapadar/index.html)

### New Python SDK Brings Native OpenTelemetry to GenAI Observability

May 15, 2025 [Alex Forbes-Reed](/content/blog/author/alex-forbes-reed/index.html)

### April Product Recap: Selene Integration, Eval Wizard Upgrades, Prompt Studio & More

May 5, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### LLM Monitoring & Evaluation for Real-World Production Use

May 5, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Systematically Improving RAG Agents

April 24, 2025 [Tahmid Tapadar](/content/blog/author/tahmid-tapadar/index.html)

### Introducing the Evaluations Wizard: How to evaluate your LLM: Building an LLM evaluation framework that actually works

April 22, 2025 [Rogerio](/content/blog/author/rogerio-chaves/index.html)

### Function Calling vs. MCP: Why You Need Both - and How LangWatch Makes It Click

April 18, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Why LLM Observability is Now Table Stakes

April 18, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### LangWatch vs. LangSmith vs. Braintrust vs. Langfuse: Choosing the Best LLM Evaluation & Monitoring Tool in 2025

April 17, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Introducing Scenario: Use an Agent to Test Your Agent

April 8, 2025 [Rogerio Chaves](/content/blog/author/rogerio-chaves/index.html)

### Tackling LLM Hallucinations with LangWatch: Why Monitoring and Evaluation Matter

April 4, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### LLM evaluations at Swis for Dutch government projects by LangWatch

April 3, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Why Your AI Team Needs an AI PM (Quality) Lead

April 2, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### LangWatch and adesso join forces: Accelerating Secure LLM Adoption for Enterprises

March 27, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### LLMOps Is Still About People: How to Build AI Teams That Don’t Implode

March 25, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Practical LLM Evaluation Framework for AI Development Teams

March 20, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### What is Model Context Protocol (MCP)? And how's LangWatch involved?

March 16, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### How PHWL.ai uses LLM Observability and Optimization to Improve AI Coaching with LangWatch

March 14, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### LangWatch.ai - Announcing - €1M funding round to bring the power of Evaluations and Auto-Optimizations to AI teams.

February 25, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### OpenAI, Anthropic, Deepseek and other LLM Providers keep dropping prices: Should you host your own model?

February 20, 2025 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### 7 Predictions for AI in 2025: A CTO's, Rogerio Chaves Perspective

January 1, 2025 [Rogerio](/content/blog/author/rogerio-chaves/index.html)

### Customer Stories: HolidayHero AI start-up <> LangWatch

December 20, 2024 CEO of HolidayHero - redated by Manouk

### LangWatch Optimization Studio - Built for AI Engineers, by AI Engineers

December 10, 2024 [Rogerio](/content/blog/author/rogerio-chaves/index.html)

### The power of MIPROv2 (DSPy) in a Low-Code environment with LangWatch’s Optimization Studio

November 10, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### What is Prompt Optimization? An Introduction to DSPy and Optimization Studio

November 7, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Deploying an OpenAI RAG Application to AWS ElasticBeanstalk

July 27, 2024 [Zhenya](/content/blog/author/zhenya/index.html)

### The complete guide for TDD with LLMs

July 3, 2024 [Rogerio - CTO](/content/blog/author/rogerio-chaves/index.html)

### Data Flywheel: Using your production data to build better LLM products

June 27, 2024 [Rogerio - CTO](/content/blog/author/rogerio-chaves/index.html)

### How Algomo reduced AI hallucinations with LangWatch

June 11, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### The AI Team: Integrating User and Domain Expert Feedback to Enhance LLM-Powered Applications

June 10, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Unit Testing Your LLM: The Power of Datasets

June 10, 2024 [Rogerio Chaves - CTO](/content/blog/author/rogerio-chaves/index.html)

### Introducing DSPy Visualizer

June 3, 2024 [Rogerio - CTO](/content/blog/author/rogerio-chaves/index.html)

### New Dutch Startup, LangWatch, brings much-needed quality control to GenAI

May 20, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### How to build a RAG application from scratch with the least possible AI Hallucinations

May 14, 2024 [Zhenya](/content/blog/author/zhenya/index.html)

### LLM Reliability with Retrieval-Augmented Generation

May 13, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Safeguarding Your First LLM-Powered Innovation: Essential Practices for Security

May 13, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### What is User Analytics for LLMs, The Difference With Traditional Analytics, And Why is it Important?

May 10, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Unlocking the Potential of Large Language Models: The LLM's Beyond the Hype

May 8, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### The 8 Types of LLM Hallucinations

May 6, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### 5 Things You Must Consider Before Putting Your Chatbot Live in Production

May 1, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Navigating the Complexities of AI-Powered Products

May 1, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Understanding Hallucinations: What are they?

April 29, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Mastering the GenAI Wave: Strategies for Success in AI Adoption

April 18, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Successfully building an AI Startup in the current booming industry

April 18, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### How Struck.build improved AI Performance with LangWatch

April 17, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Journey Through Innovation: The LLM Adventure

April 8, 2024 [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)

### Webinar recap: LLM Evaluations: Best Practices, LLM Eval types & real-world insights

Date coming soon [Manouk Draisma](/content/blog/author/manouk-draisma/index.html)
