blog

The LangWatch Blog

Engineering deep-dives, product updates, and field lessons on evaluating, testing, and observing AI agents in production.

The PR Hound: how we fixed PR review assignment with a daily agent

We leaned into agentic coding and started shipping far more code than our review process could keep up with, so PRs…

Andrew Joia · July 9, 2026

Claude vs Codex: which is the better background agent?

Some of our preferences and stories on Claude and Codex at LangWatch.

Rogerio Chaves · July 8, 2026

Background Agents on Slack: How we built our own Claude Tag before it was cool

A fleet of background agents runs a chunk of our engineering, each one scoped to one job and living in its own Slack…

Rogerio Chaves · July 5, 2026

EU AI Act compliance: are you affected?

The high-risk deadlines just moved to December 2027, but the transparency rules still land on 2 August 2026. What…

Rogerio Chaves · July 5, 2026

Cost per successful task: the metric that will decide your AI stack when model subsidies end

Cost per successful task is run cost divided by success rate. As model subsidies end, it becomes the number that…

Manouk Draisma · July 3, 2026

Introducing: Testing voice agents like you test your chat agents

Test voice agents the way you test text agents - simulated callers, traces, playback, and judge-based evaluation -…

Manouk Draisma · June 2, 2026

The Whole Platform Is Now Open Source: LangWatch May 2026 Update

May 31, 2026 Manouk Draisma

What happens when two engineering teams just... talk

May 19, 2026 Manouk Draisma

LangWatch v3.0 and the April 2026 Product Drop

April 30, 2026 Manouk Draisma

Eat Sleep Append Repeat…

April 20, 2026 Alex Forbes-Reed

Four Refactors and a Funeral: Migrating a Live System to Event Sourcing

April 20, 2026 Alex Forbes-Reed

Internal Product vs Internalised Trauma: Supporting Event Sourced Systems

April 20, 2026 Alex Forbes-Reed

Every way your AI agent can be broken (and how attackers actually do it)

April 15, 2026 Aryan

Why AI Red teaming is broken (and how we fixed it)

April 14, 2026 Rogerio Chaves

How we test Agent Skills with Scenario simulations

March 27, 2026 Sergio Cardenas

Getting to value with LangWatch, faster than ever - how to migrate from Langfuse to LangWatch with Skills.

March 26, 2026 Manouk Draisma

A Note on the LiteLLM Vulnerability

March 25, 2026 Rogerio

Product Managers and leaders are running agent simulations now, and it changing how AI ships

March 25, 2026 Sergio Cardenas

Making your AI Agent reliable: Adding Evaluations to your multi-modal agent with LangWatch Skills

March 24, 2026 Sergio Cardenas

LangWatch Skills: Your coding agent already knows how to test your agent

March 23, 2026 Manouk Draisma

Introducing LangWatch MCP: Test and evaluate AI Agents without leaving your workflow

March 12, 2026 Manouk Draisma

The Agent Development Lifecycle: Why shipping is the easy part

March 6, 2026 Manouk Draisma

The LangWatch February Drop: Cheaper Events, Claude Code, and Multimodal Evals

February 28, 2026 Manouk Draisma

New Pricing: AI growth shouldn’t increase your bill

February 20, 2026 Manouk Draisma

What is LLM monitoring? (Quality, cost, latency, and drift in production)

February 10, 2026 Manouk Draisma

What is Prompt Management? And how to version, control & deploy prompts in productions

February 10, 2026 Manouk Draisma

How OpenClaw / ClawBot works behind the scenes - and why agent observability matters

February 3, 2026 Rogerio Chaves

Instrumenting Your OpenClaw Agent with LangWatch via OpenTelemetry

February 3, 2026 Rogerio Chaves

How to Use Clawdbot + LangWatch to Monitor Your Agents in Production

February 3, 2026 Rogerio Chaves

LLM Evaluations Explained: Experiments, Online Evaluations, Guardrails, and when to use each in 2026

February 2, 2026 Rogerio Chaves

The LangWatch Monthly Drop: January 2026

January 31, 2026 Manouk Draisma

4 best tools for monitoring LLM & agent applications in 2026

January 30, 2026 Bram P

Arize AI alternatives: Top 5 Arize competitors compared (2026)

January 30, 2026 Bram P

Top 10 LLM Observability Tools: Complete Guide for 2026

January 30, 2026 Bram P

Top 5 AI evaluation tools for AI agents & products in production (2026)

January 30, 2026 Manouk Draisma

How to test AI Agents with LangWatch & Mastra / Google ADK and ship them reliably

January 29, 2026 Sergio Cardenas

Top Tools for Evaluating Voice Agents in 2025

December 30, 2025 Bram P

What are the AI Agent Events in 2026: The must-attend conferences for Agentic AI Builders

December 29, 2025 Manouk Draisma

Closing the year Strong: December Product Updates

December 24, 2025 Manouk Draisma

How to do Tracing, Evaluation, and Observability for Google ADK

December 23, 2025 Manouk Draisma

Top 5 AI Prompt Management Tools of 2025

December 23, 2025 Manouk Draisma

Writing Effective AI Evaluations, that hold up in production

December 23, 2025 Manouk Draisma

Why Agentic AI needs a new layer of testing

December 12, 2025 Manouk Draisma

Launch Week Day 5: Better Agents CLI: The reliability layer for the next wave of agent development

November 26, 2025 Rogerio Chaves

Scenario MCP: Automatic Agent Test Generation inside your editor

November 25, 2025 Aryan

Testing Voice Agents with LangWatch Scenario in Real Time

November 24, 2025 Andrew Joia

A Systematic way of Testing of AI Agents

November 20, 2025 Manouk Draisma

Introducing: LangWatch newest Prompt Playground

November 20, 2025 Andrew Garde Joia

How LangWatch helps enterprises test, evaluate, and trust their AI before release

October 27, 2025 Manouk Draisma & FlagSmith

Build vs Buy - Should you build your own LLMOps stack or leverage a purpose-built platform designed for enterprise scale?

October 17, 2025 Manouk Draisma

The 5 Best LLM Evaluation Platforms in 2025: Why LangWatch redefines the category with Agent Testing (with Simulations)

October 17, 2025 Manouk Draisma

Need-based Context Engineering: Let tests tell you what your AI agent actually needs

October 15, 2025 Andrew Joia

The Ultimate RAG Blueprint: Everything you need to know about RAG in 2025/2026

October 6, 2025 Rogerio

From Scenario to Finished: How to Test AI Agents with Domain-Driven TDD

September 26, 2025 Andrew Joia

Building Reliable AI Applications: Why Evals (and Scenarios) Are the backbone of trustworthy AI

September 25, 2025 Manouk Draisma

Are evals dead?

September 7, 2025 Rogerio Chaves

Essential LLM evaluation metrics for AI quality control: From error analysis to binary checks

September 3, 2025 Rogerio Chaves

Trace IDs in AI: LLM Observability and Distributed Tracing

August 22, 2025 Manouk Draisma

The 6 context engineering challenges stopping AI from scaling in production

August 19, 2025 Manouk Draisma

LLMOps is the new DevOps, here’s what every developer must know

August 18, 2025 Manouk Draisma

LLM observability: What is it and why it matters

August 14, 2025 Manouk Draisma

GPT-5 Release: From Benchmarks to production reality

August 8, 2025 Manouk Draisma

LLM-as-a-Judge: Using the Panel of Judges Approach to Approximate Human Preference

August 7, 2025 Rogerio Chaves

Observability Framework Design for LLM Apps - The Complete LangWatch Guide

August 1, 2025 Manouk Draisma

Top 4 Humanloop Alternatives in 2025

July 18, 2025 Manouk Draisma

Why Agent Simulations are the new Unit Tests for AI

July 7, 2025 Tahmid - AI researcher @ LangWatch

Real-time simulation visualization and debug mode

June 27, 2025 Rogerio Chaves

Scripted simulations, evaluations, and guardrails

June 26, 2025 Rogerio Chaves

Customer Story: How Roojoom automates AI Agent Quality Control with LangWatch Scenario

June 25, 2025 Manouk Draisma

Test agents on Mastra, Agno, and 10+ other frameworks

June 25, 2025 Rogerio Chaves

Introducing simulation-based agent testing

June 24, 2025 Rogerio Chaves

Why LangWatch Scenarios represents the future of AI agent testing

June 24, 2025 Rogerio Chaves

Best AI Agent Frameworks in 2025: Comparing LangGraph, DSPy, CrewAI, Agno, and More

June 21, 2025 Rogerio Chaves

Multilingual AI Agent Testing: Using Scenario to Simulate, Break, and Improve LLMs

June 20, 2025 Andrew - Engineer @ LangWatch

LangSmith Alternatives: What to use if you need more security and control

June 18, 2025 Manouk Draisma

Intro to Scenario (Testing AI agents)

June 13, 2025 Tahmid AI researcher @ LangWatch

Simulations from First Principles (How to test your agents)

June 12, 2025 Tahmid AI researcher @ LangWatch

Agent Evaluation: Framework for Testing AI Agents

June 11, 2025 Tahmid AI researcher @ LangWatch

Simulation Based Eval Framework

June 6, 2025 Tahmid - AI research @LangWatch

Introduction: The Real Issue isn’t RL

May 30, 2025 Tahmid AI Researcher @ LangWatch

Simulations to Test My Agent

May 28, 2025 Tahmid, AI Researcher @ LangWatch

New Python SDK Brings Native OpenTelemetry to GenAI Observability

May 15, 2025 Alex Forbes-Reed

April Product Recap: Selene Integration, Eval Wizard Upgrades, Prompt Studio & More

May 5, 2025 Manouk Draisma

LLM Monitoring & Evaluation for Real-World Production Use

May 5, 2025 Manouk Draisma

Systematically Improving RAG Agents

April 24, 2025 Tahmid Tapadar

Introducing the Evaluations Wizard: How to evaluate your LLM: Building an LLM evaluation framework that actually works

April 22, 2025 Rogerio

Function Calling vs. MCP: Why You Need Both - and How LangWatch Makes It Click

April 18, 2025 Manouk Draisma

Why LLM Observability is Now Table Stakes

April 18, 2025 Manouk Draisma

LangWatch vs. LangSmith vs. Braintrust vs. Langfuse: Choosing the Best LLM Evaluation & Monitoring Tool in 2025

April 17, 2025 Manouk Draisma

Introducing Scenario: Use an Agent to Test Your Agent

April 8, 2025 Rogerio Chaves

Tackling LLM Hallucinations with LangWatch: Why Monitoring and Evaluation Matter

April 4, 2025 Manouk Draisma

LLM evaluations at Swis for Dutch government projects by LangWatch

April 3, 2025 Manouk Draisma

Why Your AI Team Needs an AI PM (Quality) Lead

April 2, 2025 Manouk Draisma

LangWatch and adesso join forces: Accelerating Secure LLM Adoption for Enterprises

March 27, 2025 Manouk Draisma

LLMOps Is Still About People: How to Build AI Teams That Don’t Implode

March 25, 2025 Manouk Draisma

Practical LLM Evaluation Framework for AI Development Teams

March 20, 2025 Manouk Draisma

What is Model Context Protocol (MCP)? And how's LangWatch involved?

March 16, 2025 Manouk Draisma

How PHWL.ai uses LLM Observability and Optimization to Improve AI Coaching with LangWatch

March 14, 2025 Manouk Draisma

LangWatch.ai - Announcing - €1M funding round to bring the power of Evaluations and Auto-Optimizations to AI teams.

February 25, 2025 Manouk Draisma

OpenAI, Anthropic, Deepseek and other LLM Providers keep dropping prices: Should you host your own model?

February 20, 2025 Manouk Draisma

7 Predictions for AI in 2025: A CTO's, Rogerio Chaves Perspective

January 1, 2025 Rogerio

Customer Stories: HolidayHero AI start-up <> LangWatch

December 20, 2024 CEO of HolidayHero - redated by Manouk

LangWatch Optimization Studio - Built for AI Engineers, by AI Engineers

December 10, 2024 Rogerio

The power of MIPROv2 (DSPy) in a Low-Code environment with LangWatch’s Optimization Studio

November 10, 2024 Manouk Draisma

What is Prompt Optimization? An Introduction to DSPy and Optimization Studio

November 7, 2024 Manouk Draisma

Deploying an OpenAI RAG Application to AWS ElasticBeanstalk

July 27, 2024 Zhenya

The complete guide for TDD with LLMs

July 3, 2024 Rogerio - CTO

Data Flywheel: Using your production data to build better LLM products

June 27, 2024 Rogerio - CTO

How Algomo reduced AI hallucinations with LangWatch

June 11, 2024 Manouk Draisma

The AI Team: Integrating User and Domain Expert Feedback to Enhance LLM-Powered Applications

June 10, 2024 Manouk Draisma

Unit Testing Your LLM: The Power of Datasets

June 10, 2024 Rogerio Chaves - CTO

Introducing DSPy Visualizer

June 3, 2024 Rogerio - CTO

New Dutch Startup, LangWatch, brings much-needed quality control to GenAI

May 20, 2024 Manouk Draisma

How to build a RAG application from scratch with the least possible AI Hallucinations

May 14, 2024 Zhenya

LLM Reliability with Retrieval-Augmented Generation

May 13, 2024 Manouk Draisma

Safeguarding Your First LLM-Powered Innovation: Essential Practices for Security

May 13, 2024 Manouk Draisma

What is User Analytics for LLMs, The Difference With Traditional Analytics, And Why is it Important?

May 10, 2024 Manouk Draisma

Unlocking the Potential of Large Language Models: The LLM's Beyond the Hype

May 8, 2024 Manouk Draisma

The 8 Types of LLM Hallucinations

May 6, 2024 Manouk Draisma

5 Things You Must Consider Before Putting Your Chatbot Live in Production

May 1, 2024 Manouk Draisma

Navigating the Complexities of AI-Powered Products

May 1, 2024 Manouk Draisma

Understanding Hallucinations: What are they?

April 29, 2024 Manouk Draisma

Mastering the GenAI Wave: Strategies for Success in AI Adoption

April 18, 2024 Manouk Draisma

Successfully building an AI Startup in the current booming industry

April 18, 2024 Manouk Draisma

How Struck.build improved AI Performance with LangWatch

April 17, 2024 Manouk Draisma

Journey Through Innovation: The LLM Adventure

April 8, 2024 Manouk Draisma

Webinar recap: LLM Evaluations: Best Practices, LLM Eval types & real-world insights

Date coming soon Manouk Draisma