|
Rishi Upadhyay
I am a 5th year PhD student at UCLA where I am advised by Prof. Achuta Kadambi. I am also fortunate to collaborate with Prof. Ricky Savjani and Prof. Yuzhang Li.
My research interests are in studying and improving the physical understanding of vision diffusion models. I also work on applying frontier vision techniques and physics-informed neural networks to scientific applications in medicine and battery management.
I am honored to be awarded the 2025 Amazon Fellowship and the 2026 UCLA Jonsson Comprehensive Cancer Center Fellowship.
I attended UC Berkeley for my undergrad where I was fortunate to work with Prof. Ren Ng and Prof. Avideh Zakhor.
Email  / 
Google Scholar  / 
CV  / 
Github
I am on the job market for industry/postdoc research roles starting Spring 2027, focused on video-based world models, AI for Science, and world models for robotics.
Please reach out if you have any relevant opportunities.
|
|
|
Publications
|
Experience
|
Teaching
|
Representative/Recent papers are highlighted |
|
|
WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts
Rishi Upadhyay,
Howard Zhang,
Jim Solomon,
Ayush Agrawal,
Yunhao Ba,
Alex Wong,
Celso De Melo,
Achuta Kadambi
arXiv, 2026
PDF  / 
project page  / 
dataset
We evaluate how closely world models align with the physical world using a set of synthetic videos which test different physics concepts.
|
|
|
Battery Smart Sensing Via a Virtual Reference Electrode
Jiayi Yu*,
Aki Takahashi*,
Min-Ho Kim*,
Rishi Upadhyay*,
Howard Zhang*,
Tian-Yu Wang*,
Huayang Zhu,
Robert Kee,
Tyrone Vincent,
Bo Liu,
Xintong Yuan,
Thomas Hymel,
Kaiyan Liang,
Keyue Liang,
Haoyang Wu,
Dingyi Zhao,
Jung Tae Kim,
Achuta Kadambi,
Yuzhang Li
Nature Communications, 2026
PDF
We train a machine learning model to estimate internal electrode potentials from standard battery operating data, preventing lithium plating and enabling safer fast-charging.
Personally led the development of ML model from data to deployment.
|
|
|
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
Changwoo Jeon,
Rishi Upadhyay,
Achuta Kadambi,
ECCV, 2026
PDF  / 
project page  / 
code
Predicting pixel-aligned 3D bounding boxes from a single image without requiring camera intrinsics by using dense corner heatmaps and depth maps.
|
|
|
Learning Respiratory Dynamics in Fast Helical Free Breathing CT Imaging for Radiotherapy Planning
Rishi Upadhyay,
Pascal Paysan,
Supratik Bose,
William Delery,
Yunzheng Zhu,
William Hsu,
Daniel Low,
Stefan Schieb,
Achuta Kadambi,
Ricky R. Savjani,
Under Review, 2025
PDF  / 
We develop VAE and Diffusion based generative models that model breathing motion in lung CT scans for better radiation targeting.
|
|
|
SparseGS: Real-Time 360° Sparse View Synthesis using Gaussian Splatting
Haolin Xiong*,
Sairisheek Muttukuru*,
Rishi Upadhyay,
Pradyumna Chari,
Achuta Kadambi
3DV, 2025
PDF  / 
project page
We integrate 3D Gaussian Splatting with depth regularization and diffusion priors to enable 360° novel view synthesis from just 12 views.
|
|
|
WeatherProof: A Paired-Dataset Approach to Semantic Segmentation in Adverse Weather
Blake Gella*,
Howard Zhang*,
Rishi Upadhyay,
Tiffany Chang,
Matt Waliman,
Yunhao Ba,
Alex Wong,
Achuta Kadambi
arXiv, 2023
PDF  / 
project page
We introduce the first paired semantic segmentation dataset in adverse conditions and show that techniques that leverage this paired nature can significantly improve performance.
|
|
|
Enhancing Diffusion Models with 3D Perspective Geometry Constraints
Rishi Upadhyay,
Howard Zhang,
Yunhao Ba,
Ethan Yang,
Blake Gella,
Sicheng Jiang,
Alex Wong,
Achuta Kadambi
ACM TOG, SIGGRAPH Asia, 2023
PDF  / 
Code  / 
Project Page
We propose a loss function for latent diffusion models that improves the perspective accuracy of generated images, allowing us to create synthetic data that helps improve SOTA monocular depth estimation models.
|
|
|
Characterizing Cone Spectral Classification by Optoretinography
Vimal Prabhu Pandiyan,
Sierra Schleufer,
Emily Slezak,
James Fong,
Rishi Upadhyay,
Austin Roorda,
Ren Ng,
Ramkumar Sabesan
Biomed. Opt. Express, 2022
PDF
We characterize cone classification by ORG. Personally worked on the Retina Map Alignment Tool.
|
|
|
Scaling Building Energy Audits through Machine Learning Methods on Novel Drone Image Data
Reshma Singh,
Samuel Fernandes,
Anand Krishnan Prakash,
Paul Mathew,
Jessica Granderson,
Colman Snaith,
Rahul Pusapati,
Prakash Jadhav,
Avideh Zakhor,
Rishi Upadhyay,
Ozgur Gonen,
Harry Bergmann
Technical Report, Lawrence Berkeley National Labratory, 2022
PDF
We create an automated tool to estimate building energy efficiency through drone and thermal imaging.
|
|
|
Indoor 3D Interactive Asset Detection Using a Smartphone
Revekka Kostoeva, Rishi Upadhyay, Yersultan Sapar, Avideh Zakhor
ISPRS, 2019
PDF
We use a smartphone, AR techniques, and object classifiers to create a system to automatically detect and map objects of interest in a 3D scene.
|
|
|
Research Scientist Intern
AMD Research & Development
Jan 2026 - Aug. 2026
- Developed 3D native object-centric world model for multi-object dynamics prediction from a single image. Model disentangles simulation and rendering, allowing for fast and accurate simulation through a lightweight physics-conditioned transformer while maintaining high quality and diverse output videos through pre-trained diffusion models.
- Demonstrated that model architecture allows it to accurately model physical laws and can provide benefits for downstream robotics tasks through both synthetic data generation and in-the-loop planning on PushT and LanguageTable.
- Ported SoTA World Action Models (e.g. Cosmos 3, DreamZero, X-WAM) to AMD hardware to enable both inference and training.
|
|
|
Research Scientist Intern
Meta Reality Labs
June 2024 - Sep. 2024
- Developed end-to-end differentiable pipeline that combined deep learning and geometric eye tracking algorithms by creating a module to back-project an estimated 2D pupil ellipse into a 3D pupil.
- Demonstrated that the new pipeline was performant as a low-cost eye tracker and could be plugged into existing deep learning pipelines to improve performance across a variety of tasks.
|
|
|
Software Engineer Intern
Apple
Summer 2021
Built a custom object detection pipeline designed to run in under 5 seconds on ~40 megapixel images.
|
|
|
Software Engineer Intern
Lawrence Livermore National Laboratory
Summer 2020
Developed a tool to transfer vector fields such as force or pressure between 3D models with different shape, size, or resolution, including a fast 3D mesh parser.
|
| Course |
School |
Semester/Quarter |
Title |
| EE 102, Signals and Systems |
UCLA |
Spring 2025 |
TA |
| EE 149, Foundations of Computer Vision |
UCLA |
Fall 2024 |
TA |
| EE 149, Foundations of Computer Vision |
UCLA |
Winter 2024 |
TA |
| EE 239AS, Computational Imaging |
UCLA |
Fall 2023 |
TA |
| CS 184, Computer Graphics & Imaging |
UC Berkeley |
Spring 2022 |
Head TA |
| CS 184, Computer Graphics & Imaging |
UC Berkeley |
Spring 2021 |
Head TA |
| CS 170, Efficient Algorithms & Intractable Problems |
UC Berkeley |
Fall 2020 |
Reader |
| CS 184, Computer Graphics & Imaging |
UC Berkeley |
Summer 2020 |
TA |
| CS 184, Computer Graphics & Imaging |
UC Berkeley |
Spring 2020 |
TA |
|