DeepEval: Open-Source LLM Evaluation Framework for Python & TypeScript
The native DeepEval platform
Made by the creators of DeepEval, Confident AI is designed to scale your DeepEval AI testing workflows organization-wide with observability and collaboration.
Bye bye CSVs. Hello collaboration.
Shared evaluation dashboards
Every DeepEval test run is automatically synced to a shared dashboard. No more exporting CSVs or pasting results in Slack — your whole team sees the same metrics in real time.
Comment & annotate results
Leave comments on individual test cases and evaluation runs. Tag teammates, flag regressions, and resolve issues without switching tools.
Version datasets
Every dataset change is tracked with full version history. Roll back bad edits, compare test case coverage across versions, and know exactly what changed between evaluation runs.
Align metrics with humans
Compare metric scores against human annotations to surface false positives and negatives. Know exactly where your evals agree with your team — and where they don't.
Regression testing
Catch quality drops before they ship. Automatically compare new runs against your last known-good baseline and surface the exact test cases that regressed.
The security posture your compliance team wants.
HIPAA, SOCII COMPLIANT
Our compliance standards meets the requirements of even the most regulated healthcare, insurance, and financial industries.
MULTI-DATA RESIDENCY
Store and process data in the United States of America (North Carolina) or the European Union (Frankfurt).
RBAC AND DATA MASKING
Our flexible infrastructure allows data separation between projects, custom permissions control, and masking for LLM traces.
99.9% UPTIME SLA
We offer enterprise-level guarantees for our services to ensure mission critical workflows are always accessible.
ON-PREM HOSTING
Optionally deploy Confident AI in your cloud premises, may it be AWS, Azure, or GCP, with tailored hands-on support.
Stay in your stack. We'll meet you there.
SDKs in Python, Typescript; 20+ integrations, including OpenAI, LangGraph, Opentelemetry, and tons of more LLM gateways.
Get started today.
How is DeepEval different from Confident AI?
DeepEval is an open-source evaluation framework that lets you write and run LLM evaluation tests locally in Python or TypeScript. Confident AI is the cloud platform built on top of DeepEval that adds centralized test management, observability, collaboration, and analytics so teams can scale their evaluation workflows organization-wide.
Did Confident AI create DeepEval?
Yes. The team behind Confident AI created and maintains DeepEval. DeepEval was open-sourced to give the community a best-in-class LLM evaluation framework, while Confident AI extends it with the enterprise features teams need to operationalize evaluations at scale.
Do I need DeepEval to use Confident AI?
No. Confident AI is a standalone platform whose APIs are integrated into DeepEval. However, Confident AI is also a full LLM observability platform, so you can use it to trace, monitor, and evaluate your LLM applications in one place — no more siloing evals and tracing across different tools.
Is DeepEval free to use?
Yes. DeepEval is fully open-source under the Apache 2.0 license and free to use for any purpose. Confident AI offers a free tier as well, along with paid plans for teams that need advanced features like role-based access, custom dashboards, and dedicated support.
What evaluation metrics does DeepEval support?
DeepEval ships with 50+ research-backed metrics including faithfulness, answer relevancy, contextual recall, contextual precision, hallucination, bias, toxicity, and more. You can also define fully custom metrics using Python or LLM-as-a-judge approaches.
Can I use Confident AI with other evaluation frameworks?
Yes. While Confident AI has first-class support for DeepEval, it also integrates with other popular tools and frameworks through its REST API and SDKs, so you can centralize results regardless of how you run your evaluations.