DEV Community

#evaluation

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Your Agent Said It Worked. Go Check the World, Not the Sentence.

Your Agent Said It Worked. Go Check the World, Not the Sentence.

Comments 1
4 min read
How EvalPort's Grader System Works: 11 Types for LLM Evaluation

How EvalPort's Grader System Works: 11 Types for LLM Evaluation

Comments
2 min read
Your New Eval Rule Is Untested Code Guarding Production

Your New Eval Rule Is Untested Code Guarding Production

Comments 1
5 min read
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

Comments 2
7 min read
OpenEval: Why LLM Evaluation Needs a Standard Format

OpenEval: Why LLM Evaluation Needs a Standard Format

Comments
1 min read
Right Tool, Wrong Arguments: The Agent Failure Your Evals Wave Through

Right Tool, Wrong Arguments: The Agent Failure Your Evals Wave Through

4
Comments 2
4 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

1
Comments
6 min read
Your Agent's Confidence Score Is Not a Probability

Your Agent's Confidence Score Is Not a Probability

3
Comments 1
4 min read
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs

PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs

Comments
3 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Comments
7 min read
Your Agent's Deadline Is a Correctness Test, Not an SLO

Your Agent's Deadline Is a Correctness Test, Not an SLO

3
Comments 1
4 min read
Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One

Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One

2
Comments 3
5 min read
Evaluating LLM Apps in Python

Evaluating LLM Apps in Python

Comments
9 min read
Evaluating LLM Apps in Java

Evaluating LLM Apps in Java

Comments
10 min read
Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference

Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference

1
Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.