Projects

Research, projects, and writing.

  1. AutoHarnessBench landing page and overall leaderboard
    AutoHarnessBench

    2026, Project

    A long-horizon SWE & self-improvement benchmark evaluating automatic harness optimization.

  2. Privileged self-distillation: rollout, construct hint, distill
    Privileged Self-Distillation

    2026, Research

    A post-training method that distills training signal from failed rollouts corrected with privileged hints.

  3. Meta-Reward: evaluator harness optimized by meta-reward, feeding a frozen LLM judge
    Meta-Reward: Reward Modeling as Harness Optimization

    2026, Research

    Framing reward modeling as a harness optimization problem.

  4. 67%
    80%
    87%
    Meta-Agent: Continual Learning for Agents

    2026, Project

    An open-source framework for agent continual learning.

  5. Canvas Coworker — AI coworker for knowledge work
    Canvas Coworker

    2026, Project

    An AI coworker for knowledge work.

  6. DashGPT landing page — Data to Dashboard magic
    DashGPT

    2024, Project

    An AI that turns spreadsheets into interactive dashboards instantly. No code, just insights.

  7. Goldfish long-video retrieval framework architecture
    Goldfish: Vision-Language Understanding of Arbitrarily Long Videos

    ECCV 2024, Research

    A SOTA vision-language model for question answering on hour-long videos. Harvard, contributor.

  8. MiniGPT4-Video architecture: interleaved visual-textual tokens through ViT and LLM
    MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding

    CVPR Workshop 2024, Research

    A multimodal LLM that interleaves visual and textual tokens for video understanding. Harvard, contributor.

  9. SlowFormer energy attack on vision transformers
    SlowFormer: Adversarial Robustness for Efficient Vision Transformers

    CVPR 2024, Research

    A universal patch attack that maximizes the compute cost of efficient vision transformers. UC Davis, with K L Navaneet et al.

  10. Do Function Vectors Factor Task and Distribution?

    Preprint 2024, Research

    An interpretability study of how LLMs represent tasks during in-context learning. Harvard & MIT, with Jacob Andreas.

  11. A human and an LLM choose different quantifiers for the same quantity
    Can LLMs Understand Quantifiers Like Humans?

    2023, Research

    Uses parallel Rational Speech Act models to compare human and GPT-3.5/GPT-4 quantifier choices across two context conditions. Harvard & MIT, with guidance from Ced Zhang and Joshua Tenenbaum.

  12. Continual self-supervised learning with a frozen teacher preserving old representation geometry
    Continual Self-Supervised Learning with Knowledge Distillation

    Preprint 2023, Research

    A video embedding model for Twitch that learns new content without forgetting the old. Amazon Science, with Xiangbo Li and Saad Ali.

  13. Mitigating Negative Transfer in Multi-Task Learning with EMA Loss Weighting

    AAAI 2023, Research

    Student abstract on EMA loss weighting strategies that reduce negative transfer in multi-task learning. Stanford.

  14. Deep Learning-Based Autism Spectrum Disorder Detection Using Emotion Features From Video

    JMIR Biomedical Engineering 2022, Research

    Detecting autism spectrum disorder from emotion features in video recordings. Stanford.

  15. TikTokFER identity resolution across videos and duplicate removal
    TikTok for Good: Creating a Diverse Emotion Expression Database

    CVPR Workshops 2022, Research

    A diverse emotion expression database collected from TikTok. Stanford, contributor.

  16. Komma group-meeting scheduling landing page
    Komma

    2020–2022, Product

    A group-meeting scheduling product for sales teams. Co-founder.