My research sits at the intersection of computer graphics, vision, machine learning, and program synthesis. I study visual programs, symbolic procedures whose executions produce or analyze shapes, scenes, or motion. I investigate two complementary directions: how learning systems can author visual programs in a given language, and how to adapt the language itself for a particular task.
My research sits at the intersection of computer graphics, vision, machine learning, and program synthesis. AI systems have transformed how we generate and analyze visual data, but they provide a poor interface for collaborative work, producing outputs that are hard to control, validate, and revise. To create useful visual content, we require a representation that respects and exposes the task-specific concepts that matter to the people and systems that use it.
My work explores how code can provide such a representation. I study visual programs, symbolic procedures whose executions produce or analyze shapes, scenes, or motion. People use code to formalize intent through executable operations, and networks can learn to model collections of programs, but the usefulness of a program depends on the operations and abstractions it can access. Therefore, I investigate two complementary directions: how learning systems can author visual programs in a given language, and how to adapt the language itself for a particular task. Toward the first, I've developed methods for inferring programs that explain visual data and creating visual content with programs. Toward the second, methods for discovering reusable visual abstractions.
My long-term aim is to make visual programs the default substrate for learning systems that model visual data. With the right underlying representation, generative AI can move beyond producing one-off artifacts and instead allow us to generate, inspect, and manipulate the visual world with the precision and flexibility we expect from code.
News
November 2026Visiting UC Berkeley for an invited talk
October 2026Visiting UCSD for an invited talk
September 2026VibeAnimation awarded a Magic Grant from the Brown Institute (with Jiaju Ma)
Topic: Discovering visual abstractions and concepts
The quality of a visual program depends on the operations its language provides, but language design is expensive and time-consuming. I develop methods that discover visual abstractions from data, learned priors, and human guidance. These abstractions simplify learning, capture meaningful variation, and provide people and machines with more useful ways to control visual content.
Topic: Creating visual content with programs
Generative models can produce impressive visual content, but their outputs are hard to control and revise. To overcome this limitation, I develop neurosymbolic methods that learn to generate and edit visual content through programs.
Topic: Inferring programs that explain visual data
Recovering a program that explains a visual observation supports reverse engineering, manipulation, and structural analysis, but visual data rarely comes with program annotations. I develop self-supervised methods that use execution, search, and iterative rewriting to train networks that learn to infer visual programs from unlabeled data.
Topic: Analyzing and verifying visual properties
Many visual properties are difficult to recognize or verify, especially when annotations and reliable specification checkers are not readily available. I develop structured analysis methods for geometric and spatial reasoning under weak or self-supervised regimes.
Preprints
ShapeLib: Designing a Library of Programmatic 3D Shape Abstractions with Large Language Models
We guide an LLM to design a library of abstractions from a small set of seed shapes and descriptions. These functions generalize to new shapes and expose semantically aligned, easy-to-work-with interfaces, supporting downstream tasks like shape editing and generation.
Self-Consistency for LLM-Based Motion Trajectory Generation and Verification
We present a self-consistency method that enables more accurate LLM-based trajectory generation without supervision and show that it can be used for trajectory verification.
Procedural Scene Programs for Open-Universe Scene Generation: LLM-Free Error Correction via Program Search
We develop a framework for part-level concept learning from single-image examples that enables text-to-image diffusion models to compose novel objects from meaningful components.
Neurosymbolic Methods for Shape Analysis and Generation
We design a system that learns how to edit visual programs. Given an initial program and a visual target, our network predicts local edit operations that can be applied to the input program to improve its similarity to the target.
We introduce ParSEL, a system that enables controllable editing of high-quality 3D assets from natural language. Given a segmented 3D mesh and an editing request, ParSEL produces a parameterized editing program that allows users to explore a family of shape variations with precise controls.
Learning to Infer Generative Template Programs for Visual Concepts
We develop a neurosymbolic method that learns how to infer Template Programs; partial programs that capture visual concepts in a domain-general fashion. Our framework supports multiple concept-related tasks: cosegmentation, few-shot generation, and concept synthesis.
Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases
We explore how code rewriting can be used to improve visual program induction networks. Across multiple domains, our family of visual program rewriters treat programs as structured objects to produce better training targets for bootstrapped learning methods.
Oral Presentation
ShapeCoder: Discovering Abstractions for Visual Programs from Unstructured Primitives
ShapeCoder automatically discovers abstraction functions, and infers visual programs that use these abstractions, to compactly explain an input dataset of shapes represented with unstructured primitives. Discovered abstractions capture common patterns (both structural and parametric) so that programs rewritten with these abstractions are more compact and expose fewer degrees of freedom.
We summarize research on neurosymbolic models in computer graphics: methods that combine the strengths of both AI and symbolic programs to represent, generate, and manipulate visual data.
SHRED: 3D Shape Region Decomposition with Learned Local Operations
SHRED is a method for 3D SHape REgion Decomposition that consumes a 3D shape as input and uses learned local operations to produce a segmentation that approximates fine-grained part instances.
PLAD: Learning to Infer Shape Programs with Pseudo-Labels and Approximate Distributions
We group a family of shape program inference methods under a single conceptual framework, where training is performed with maximum likelihood updates sourced from either Pseudo-Labels or an Approximate Distribution (PLAD). Compared with policy gradient reinforcement learning, we show that PLAD techniques infer more accurate shape programs and converge significantly faster.
The Neurally-Guided Shape Parser: Grammar-based Labeling of 3D Shape Regions with Approximate Inference
We frame 3D shape semantic segmentation as a label assignment problem over shape regions; in this paradigm, we show our approximate inference formulation improves performance over comparison methods that (i) use regions to group per-point predictions, (ii) use regions as a self-supervisory signal, or (iii) assign labels to regions under alternative formulations.
An algorithm that automatically discovers macro operators that are useful for collections of 3D shape programs. It discovers macros that make programs more compact by minimizing the number of function calls and free parameters required to represent a dataset of imperative programs that may contain continuous parameters. We show these discovered macros improve performance on down-stream tasks such as program inference from unstructured geometry, generative modeling, and goal-directed editing.
ShapeAssembly: Learning to Generate Programs for 3D Shape Structure Synthesis
Using a hybrid neural-procedural approach, we present a deep generative model that learns to synthesize 3D shapes by writing programs in ShapeAssembly, a domain-specific 'assembly language' for 3D shape structures.