DEV Community

#computervision

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Decoupling Physical Control and Reasoning: DeepMind's Gemini Robotics 2 Architecture

Decoupling Physical Control and Reasoning: DeepMind's Gemini Robotics 2 Architecture

Comments
5 min read
Segmentation Is Not the Finish Line

Segmentation Is Not the Finish Line

Comments
5 min read
The mAP50 That Lied to Me: A Debugging Story About On-Device Safety AI

The mAP50 That Lied to Me: A Debugging Story About On-Device Safety AI

Comments
5 min read
AI Landscape Design in 2026: Why the Generative Model Is the Easy Part

AI Landscape Design in 2026: Why the Generative Model Is the Easy Part

Comments
5 min read
From 42 Points to 842: Debugging a CPU-Only 3D Reconstruction Pipeline

From 42 Points to 842: Debugging a CPU-Only 3D Reconstruction Pipeline

Comments
2 min read
Detecting Objects in Satellite Imagery with YOLOv8: xView + DOTA to YOLO in Practice

Detecting Objects in Satellite Imagery with YOLOv8: xView + DOTA to YOLO in Practice

1
Comments 3
5 min read
The Model Is the Easy Part: What a Real-Time Computer Vision Product Actually Takes

The Model Is the Easy Part: What a Real-Time Computer Vision Product Actually Takes

Comments
6 min read
The Master Woodworker

The Master Woodworker

2
Comments
3 min read
How Vision-Language Models Learned to Reason About Space (10 Papers, One Thread)

How Vision-Language Models Learned to Reason About Space (10 Papers, One Thread)

Comments
3 min read
Identity Consistency Is the Hard Part of Two-Person Image Generation

Identity Consistency Is the Hard Part of Two-Person Image Generation

Comments
2 min read
Deep Learning & Computer Vision in Web Diffing: Solving Layout Shifts with Neural Embeddings and SSIM

Deep Learning & Computer Vision in Web Diffing: Solving Layout Shifts with Neural Embeddings and SSIM

Comments 1
5 min read
Bridging the Frame Gap: Robot-Centric Pointmaps for VLA Models

Bridging the Frame Gap: Robot-Centric Pointmaps for VLA Models

Comments
3 min read
Audio-Visual Flamingo: Advancing Open-Source Intelligence for Long-Form Video Reasoning

Audio-Visual Flamingo: Advancing Open-Source Intelligence for Long-Form Video Reasoning

Comments
3 min read
CarSegNet v2: One Lot Photo to a Showroom Composite

CarSegNet v2: One Lot Photo to a Showroom Composite

Comments
14 min read
A Hands-On Guide to kalbee: Your First Kalman Filter (and Beyond)

A Hands-On Guide to kalbee: Your First Kalman Filter (and Beyond)

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.