A long-horizon benchmark evaluating harness-driven agent self-improvement.
A post-training method distilling privileged corrections to failed rollouts.
Framing reward modeling as a harness optimization problem.
An open-source framework for agent continual learning.
A production coding agent for long-horizon GTM and knowledge-work tasks with 100+ integrations.
A SOTA vision-language model for question answering on hour-long videos.
Adversarial robustness for vision transformers.
A Twitch video embedding model that continually learns from video streams.
An interpretability study of how LLMs represent tasks during in-context learning.
A study comparing human and LLM quantifier choices using Rational Speech Act models.
A multimodal LLM that interleaves visual and textual tokens for video understanding.
A system that turns spreadsheets into interactive dashboards.
EMA-based loss weighting that reduces negative transfer in multi-task learning.
Detecting autism spectrum disorder from emotion features in video recordings.
A diverse emotion-expression dataset collected from TikTok videos.
A group-meeting scheduling product for sales teams. Co-founder.