Smiling Buddha Open Research
Reproducible baselines

Physical AI benchmarks.

Standardised evaluation suites for physical AI systems. Every benchmark ships with data, protocol, reference implementations, and a public leaderboard.

Active benchmarks

Suites currently open for submission

SB-Perception-01
Long-horizon object tracking in cluttered environments
Track objects through occlusion, motion blur, and lighting changes across a 45-minute continuous capture. Includes reference PyTorch baseline.
perceptiontrackingv2.1
842 submissionsOpen
SB-Manip-03
Precision insertion under partial observability
Fork-and-pallet-jack style insertion tasks under camera occlusion. Success measured on physical fidelity, tolerance, and cycle time.
manipulationmaterial handling
231 submissionsOpen
SB-Nav-02
Dynamic-environment path planning
Navigate through a busy warehouse simulated distribution centre with moving humans and vehicles. Evaluated on safety, efficiency, and stop-events.
navigationsafety
489 submissionsOpen
SB-Cleaning-01
Floor-cleaning coverage under obstacle drift
Achieve target coverage while obstacles change positions between passes. Scored on coverage, redundancy, and shine-index outcome.
cleaningplanning
128 submissionsOpen
SB-Safety-01
Edge-case classification under distribution shift
Classify novel edge cases correctly across held-out distribution shifts. Scored on precision, recall, and false-safe rate.
safetyclassification
67 submissionsBeta
SB-Sim2Real-01
Sim-to-real transfer for locomotion primitives
Train in simulation, deploy on real hardware. Evaluated on transfer gap, training-data efficiency, and reproducibility.
sim-to-reallocomotion
Coming soon
How submission works

Three-step protocol

01
Download the benchmark bundle
Each benchmark ships with data, evaluation code, reference baseline, and reproducibility protocol. All MIT-licensed.
02
Run our reference protocol
Follow the documented protocol so your results are comparable to every other submission. No hyperparameter tuning post-hoc.
03
Submit your results + code
Push to the benchmark's GitHub repo. Community-reviewed; verified submissions appear on the public leaderboard.