FAMOSFeed-Forward 3D Articulation Modeling
from Sparse Observations

  1. Kevin Qu1,2
  2. Tao Sun1
  3. Massimiliano Viola1
  4. Liyuan Zhu1
  5. Zhizhuo Zhou1
  6. Sayan Deb Sarkar1
  7. Konrad Schindler2
  8. Iro Armeni1

1 Stanford University2 ETH Zürich

Left: casual captures of a stove cabinet and a red cabinet in different articulation states are lifted to partial point clouds, from which FAMOS predicts colored movable parts and joint axes. Right: predictions on real photographs of a wall cabinet, a nightstand and a laptop.

Summary

FAMOS is a feed-forward method that predicts movable-part segmentation and joint parameters from a sparse set of monocular observations. By jointly reasoning over the whole input set, it grounds articulation prediction in observed motion rather than shape priors alone. The model can handle a variable number of inputs (including a single view) and is trained at scale by extending the training data with assets from our procedural data generator.

  • +28.2 pts Movable-part segmentation F1 gain on ACD over the strongest feed-forward baseline
  • +27.8 pts Motion estimation MAO F1 gain on ACD over the strongest feed-forward baseline
  • 0.1 s Inference time nearly 3,000× faster than per-object optimization

Abstract

Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a single observation and therefore rely heavily on learned category-level shape priors. We present FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds. Our model jointly reasons over multiple observations and naturally supports a variable number of inputs, including a single view. To aggregate articulation cues across observations, we introduce a Multi-state Articulation Transformer with alternating state-wise and global attention. We further propose an observed articulation span objective that supervises the motion range each part exhibits across the input observations, encouraging the model to leverage the full observation set. To overcome the limited scale and diversity of existing datasets, we introduce a procedural data generator that synthesizes self-annotated assets during training. Experiments on PartNet-Mobility, ACD, and ArtiCraft-10K demonstrate consistent improvements over both feed-forward and optimization-based baselines.

How It Works

Multi-State Articulation Transformer

FAMOS architecture: partial point clouds are encoded with PartField and geometric embeddings, then processed with part queries through alternating state-wise and global attention. Decoder heads predict movable-part segmentation, joint parameters, and the auxiliary observed articulation span.
  1. 01

    Point encoding

    Each observation is lifted to a partial point cloud, then tokenized with a frozen PartField backbone and geometric embeddings. No state-index embedding is added, preventing the model from learning state-index-specific identities and allows it to flexibly process a variable number of inputs during inference.

  2. 02

    Alternating attention

    Each transformer layer alternates between two attention stages. State-wise attention lets the point tokens of each input view aggregate context over the geometry of their own partial point cloud. Global attention then refines all point and query tokens jointly, allowing global information exchange across observations.

  3. 03

    Prediction heads

    Per-query heads decode joint type, axis and origin; a per-point head decodes the movable-part segmentation. Because one query segments a part across every observation, the segmentation is consistent between states.

Observed Articulation Span

Three input views of a cabinet with a drawer and a door. The observed articulation span is illustrated as 25 cm of drawer travel and 50 degrees of door rotation covered across those views.

We introduce an auxiliary objective that predicts, for each movable part, the range of joint motion observed across the input states. Since this quantity is defined over the full observation set, it encourages consistent part localization and aggregation of motion evidence across states, rather than relying on a single observation. The span is also invariant to the unknown canonical rest pose, making it a well-defined supervision signal for sparse, unordered inputs.

Procedural Training Data

To increase the scale and diversity of our training data, we design a procedural asset generation pipeline that assembles articulated objects from basic shape primitives (boxes, cylinders, hollow cylinders, half-tori, spheres, holed panels). The procedural generation introduces minimal computational overhead and can run on-the-fly during training in the dataloader.

object families

seed

Comparison to Baselines

Cabinet with drawers

PartNet-Mobility

Two-door cabinet

PartNet-Mobility

Vanity desk

ACD

Sideboard

ACD

FAMOS visualization: detached point clusters and one unsupported axis omitted.

Stove cabinet

ArtiCraft-10K

Drawer cabinet

ArtiCraft-10K

Static basePredicted movable partsRed axes: revolute jointsYellow axes: prismatic jointsReArt receives complete point clouds. Particulate is re-trained on our data.For clarity, predictions on one input view are shown.

Real-World Objects

Despite being trained on synthetic data only, FAMOS generalizes to casual real-world captures.

RGB viewscasual phone captures VGGT-Ω reconstructionmasks from SAM2 FAMOS predictionmovable part segmentation and joint axes

Drawer unit

RGB viewsPhotograph 1 of 3 of the drawer unit.Photograph 2 of 3 of the drawer unit.Photograph 3 of 3 of the drawer unit.
VGGT-Ω reconstructionThe drawer unit as a dense colored point cloud, reconstructed from the photographs.
FAMOS predictionFAMOS prediction for the drawer unit: the static base in grey, three prismatic drawers, each with a yellow slide axis.

Laptop

RGB viewsPhotograph 1 of 2 of the laptop.Photograph 2 of 2 of the laptop.
VGGT-Ω reconstructionThe laptop as a dense colored point cloud, reconstructed from the photographs.
FAMOS predictionFAMOS prediction for the laptop: the static base in grey, a revolute lid with a red hinge axis along its bottom edge.

Metal cabinet

RGB viewsPhotograph 1 of 2 of the metal cabinet.Photograph 2 of 2 of the metal cabinet.
VGGT-Ω reconstructionThe metal cabinet as a dense colored point cloud, reconstructed from the photographs.
FAMOS predictionFAMOS prediction for the metal cabinet: the static base in grey, a prismatic drawer and a revolute door.

Nightstand

RGB viewsPhotograph 1 of 2 of the nightstand.Photograph 2 of 2 of the nightstand.
VGGT-Ω reconstructionThe nightstand as a dense colored point cloud, reconstructed from the photographs.
FAMOS predictionFAMOS prediction for the nightstand: the static base in grey, two prismatic drawers, each with a yellow slide axis.

Kitchen cabinet

RGB viewsPhotograph 1 of 2 of the kitchen cabinet.Photograph 2 of 2 of the kitchen cabinet.
VGGT-Ω reconstructionThe kitchen cabinet as a dense colored point cloud, reconstructed from the photographs.
FAMOS predictionFAMOS prediction for the kitchen cabinet: the static base in grey, a revolute door with a red hinge axis at its edge.

Cane cabinet

RGB viewsPhotograph 1 of 2 of the cane cabinet.Photograph 2 of 2 of the cane cabinet.
VGGT-Ω reconstructionThe cane cabinet as a dense colored point cloud, reconstructed from the photographs.
FAMOS predictionFAMOS prediction for the cane cabinet: the static base in grey, a revolute door with a red hinge axis.

Wall cabinet

RGB viewsPhotograph 1 of 2 of the wall cabinet.Photograph 2 of 2 of the wall cabinet.
VGGT-Ω reconstructionThe wall cabinet as a dense colored point cloud, reconstructed from the photographs.
FAMOS predictionFAMOS prediction for the wall cabinet: the static base in grey, a revolute door with a red hinge axis.

Hardcover book

RGB viewsPhotograph 1 of 2 of the hardcover book.Photograph 2 of 2 of the hardcover book.
VGGT-Ω reconstructionThe hardcover book as a dense colored point cloud, reconstructed from the photographs.
FAMOS predictionFAMOS prediction for the hardcover book: the static base in grey, a revolute cover with a red hinge axis along the spine.

Static basePredicted movable partsRed axes: revolute jointsYellow axes: prismatic joints

Quantitative Results

FAMOS consistently outperforms the baselines in movable-part segmentation and articulation estimation, for both multi-state and single-state settings.

F1 (%) · higher is better

Particulate is re-trained on our data.

Movable-part segmentation F1 at IoU ≥ 0.5. † Re-trained on our data. +n: gain over the strongest baseline.

Runtime per object

Inference Time

Four input states on one NVIDIA RTX 5090 · logarithmic scale.

Why Multiple States Matter

Articulation estimation from a single state is intrinsically ambiguous: several kinematic models explain the same visible geometry, causing the prediction to fall back on learned category-level shape priors. FAMOS instead grounds its prediction in the motion cues that additional observations reveal.

Two objects, one per row. Each row shows two RGB views, the prediction from a single state, the prediction from two states, and the ground truth. From one state the predicted joint is the wrong type or on the wrong edge; with the second state the prediction matches the ground truth.
Estimating articulation from a single input state can yield a plausible but wrong prediction. A second input state reveals the underlying motion and allows the model to predict the correct joint configuration.revoluteprismatic
ArtiCraft-10K · F1 (%)
Movable-part segmentationMotion estimation
Accuracy against the number of input states on ArtiCraft-10K Movable-part segmentation rises from 80.4 to 86.6 F1 from one to two observations. Motion estimation rises from 67.6 to 77.7. Later observations bring smaller gains, and performance remains stable beyond four inputs. 6470778490 12345678910 Number of input states

Citation

TBD — BibTeX citation coming soon.