Wonjae (Dan) Kim
김원재 · Lead Research Scientist · Multimodal AI
At TwelveLabs, I lead the Marengo & Search team, where we build video-native multimodal representations and search systems. Most recently, we launched Marengo 3.5, a multimodal embedding model designed around the temporal structure of video.
I’m a co-first author of ViLT (cited: 3040), an early work on efficient vision-language modeling. Previously, I was a research scientist at Naver AI LAB and Kakao, and I hold an M.Sc. and B.Sc. from Seoul National University.
My current research focuses on:
- Video-native Multimodal Representation Learning (video, audio, image, text, documents)
- Large-scale Embedding & Search Systems
- User Behavior Modeling for Search
I also write about information, AI tools, and coordination. Recent essays include Speak Concisely, Write Verbosely and The Theory of Consensus Reset Cost Optimization.
At TwelveLabs, we’re hiring researchers and engineers to work on multimodal embeddings and video search. Open role in Seoul → · Coffee chat
news
| Sep 01, 2026 | We launched Marengo 3.5, a video-native multimodal embedding model built around the temporal structure of video. |
|---|---|
| Apr 19, 2026 | TwelveLabs launches Pegasus 1.5, a state-of-the-art video-language model for understanding and reasoning over long-form video. |
| Dec 01, 2025 | TwelveLabs releases Marengo 3.0, a new standard for foundation models that understand the world in all its complexity. |
| Apr 01, 2025 | One CVPR-2025 EVAL-FoMo 2 Workshop paper: Emergence of Text Readability in Vision Language Models. |
| Feb 04, 2025 | I’ve started a new chapter at TwelveLabs! |
latest posts
| Feb 22, 2026 | The Theory of Consensus Reset Cost Optimization |
|---|---|
| Feb 13, 2026 | Speak Concisely, Write Verbosely |
| Jun 11, 2025 | The Gentle Singularity |
selected publications
-
- HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts(cited: 26)In 18th European Conference on Computer Vision (ECCV 2024), 2024