Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Your Agent Said It Worked. Go Check the World, Not the Sentence.
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 5
Your Agent Said It Worked. Go Check the World, Not the Sentence.
#
ai
#
agents
#
evaluation
#
observability
Comments
1
comment
4 min read
How EvalPort's Grader System Works: 11 Types for LLM Evaluation
Adha AK
Adha AK
Adha AK
Follow
Aug 4
How EvalPort's Grader System Works: 11 Types for LLM Evaluation
#
llm
#
evaluation
#
testing
#
opensource
Comments
Add Comment
2 min read
Your New Eval Rule Is Untested Code Guarding Production
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 2
Your New Eval Rule Is Untested Code Guarding Production
#
ai
#
agents
#
evaluation
#
testing
Comments
1
comment
5 min read
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother
Xinyang Wu
Xinyang Wu
Xinyang Wu
Follow
Aug 3
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother
#
rag
#
llm
#
embeddings
#
evaluation
Comments
2
comments
7 min read
OpenEval: Why LLM Evaluation Needs a Standard Format
Adha AK
Adha AK
Adha AK
Follow
Jul 30
OpenEval: Why LLM Evaluation Needs a Standard Format
#
llm
#
evaluation
#
ai
#
testing
Comments
Add Comment
1 min read
Right Tool, Wrong Arguments: The Agent Failure Your Evals Wave Through
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 26
Right Tool, Wrong Arguments: The Agent Failure Your Evals Wave Through
#
ai
#
agents
#
evaluation
#
observability
4
reactions
Comments
2
comments
4 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
Maya Andersson
Maya Andersson
Maya Andersson
Follow
Jul 21
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
#
statistics
#
machinelearning
#
datascience
#
evaluation
1
reaction
Comments
Add Comment
6 min read
Your Agent's Confidence Score Is Not a Probability
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 29
Your Agent's Confidence Score Is Not a Probability
#
ai
#
agents
#
evaluation
#
observability
3
reactions
Comments
1
comment
4 min read
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 15
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs
#
llms
#
codegeneration
#
selfrepair
#
evaluation
Comments
Add Comment
3 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.
Muhammed Rasin O M
Muhammed Rasin O M
Muhammed Rasin O M
Follow
Jul 10
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.
#
evaluation
#
dataagents
#
benchmarks
#
syntheticdata
Comments
Add Comment
7 min read
Your Agent's Deadline Is a Correctness Test, Not an SLO
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 31
Your Agent's Deadline Is a Correctness Test, Not an SLO
#
ai
#
agents
#
evaluation
#
observability
3
reactions
Comments
1
comment
4 min read
Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 19
Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One
#
ai
#
agents
#
evaluation
#
observability
2
reactions
Comments
3
comments
5 min read
Evaluating LLM Apps in Python
Puneet Gupta
Puneet Gupta
Puneet Gupta
Follow
Jul 5
Evaluating LLM Apps in Python
#
python
#
ai
#
llm
#
evaluation
Comments
Add Comment
9 min read
Evaluating LLM Apps in Java
Puneet Gupta
Puneet Gupta
Puneet Gupta
Follow
Jul 5
Evaluating LLM Apps in Java
#
java
#
ai
#
llm
#
evaluation
Comments
Add Comment
10 min read
Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 2
Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference
#
ai
#
agents
#
evaluation
#
typescript
1
reaction
Comments
Add Comment
5 min read
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account