PaperScope
LIVE · 2026-10-05 05:40 UTC

RYOPO: Bringing End-to-End Category-Level Object Pose Estimation into Real Time

Hakjin Lee, Junghoon Seo, Jaehoon Sim

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.03013 v1
Category
Submitted
2026-10-02

Abstract

Category-level object pose estimation predicts the rotation, translation, and metric size of unseen instances within known categories. Many accurate RGB-D methods rely on external instance segmentation and crop-based pose estimation, introducing separate stages and object-dependent processing costs that hinder real-time inference. To bring accurate pose estimation into real time, we present \ours{}, an end-to-end trainable query-based RGB-D set predictor. It jointly detects and segments objects and estimates their \mbox{9-DoF} poses without explicit CAD-derived shape priors or a separately trained instance segmentor. Shared image and scene encoding avoids repeated per-object crop encoding. A query-conditioned geometry pathway associates observed 3D points and RGB features with object queries and incorporates shared scene context. Object-centric refinement uses the resulting point descriptors to update an explicit pose state through pose-conditioned cross-attention and recurrent residual corrections. On NOCS, \ours{} substantially improves on published RGB-D joint detection and pose estimation results. It achieves competitive performance compared with two-stage methods under all-object evaluation on REAL275 and HouseCat6D, while enabling real-time full-frame pose estimation at $31.8$ FPS on an RTX~A6000. Project page: https://yopo-series.github.io/RYOPO-project-page/.

Comment: Project page: https://yopo-series.github.io/RYOPO-project-page/

arXiv abs page · PDF · same-day batch