Research, projects, and writing.
A long-horizon SWE & self-improvement benchmark evaluating automatic harness optimization.
A post-training method that distills training signal from failed rollouts corrected with privileged hints.
Framing reward modeling as a harness optimization problem.
An open-source framework for agent continual learning.
A SOTA vision-language model for question answering on hour-long videos. Harvard, contributor.
A multimodal LLM that interleaves visual and textual tokens for video understanding. Harvard, contributor.
A universal patch attack that maximizes the compute cost of efficient vision transformers. UC Davis, with K L Navaneet et al.
An interpretability study of how LLMs represent tasks during in-context learning. Harvard & MIT, with Jacob Andreas.
Uses parallel Rational Speech Act models to compare human and GPT-3.5/GPT-4 quantifier choices across two context conditions. Harvard & MIT, with guidance from Ced Zhang and Joshua Tenenbaum.
A video embedding model for Twitch that learns new content without forgetting the old. Amazon Science, with Xiangbo Li and Saad Ali.
Student abstract on EMA loss weighting strategies that reduce negative transfer in multi-task learning. Stanford.
Detecting autism spectrum disorder from emotion features in video recordings. Stanford.
A diverse emotion expression database collected from TikTok. Stanford, contributor.