Research Intern, UMAP, Snap Research
Bellevue, WA
Studying decoder-only LLMs as general-purpose embedding models, with recent work on information flow, retrieval quality, long-context reasoning, and sparse attention mechanisms.
- Proposed Hierarchical Token Prepending, a training-free method that improves decoder-based LLM embeddings on retrieval and general embedding benchmarks.
- Investigated RLVR-based search agents and LLM-driven query rewriting for open-domain retrieval and recommendation settings.
- Collaborated on Threshold Differential Attention, a sink-free sparse attention mechanism for long-context language models.