I am a final-year Ph.D. candidate at the Machine Learning and Intelligence Lab. at Korea Advanced Institute of Science and Technology (KAIST), advised by Prof. Eunho Yang.

Research Focus

My research centers on efficient inference for large-scale foundation models, with a broader interest in efficient training and post-training. I primarily approach efficiency from an algorithmic perspective, but I am especially interested in methods that remain effective under real hardware and system constraints. In other words, I care not only about improving algorithmic metrics, but also about whether those improvements translate into meaningful reductions in latency, memory usage, and computational cost in realistic workloads. This perspective has naturally led me to study the interaction between algorithms and the characteristics of modern accelerators, memory hierarchies, parallel execution, and serving systems.

Current Research Directions

My recent work focuses on several complementary directions toward practical acceleration of large-scale models. One major direction is efficient reinforcement learning for language models, particularly accelerating RLVR rollouts through techniques such as quantized inference and efficient scheduling. I am also continuing to work on speculative decoding, with an emphasis on making it genuinely effective in realistic serving settings such as large-batch, high-throughput, and offline inference, rather than optimizing acceptance-related metrics in isolation. Quantization is another central theme in my research, both as a standalone compression technique and as a building block that can be integrated with other acceleration methods. More broadly, I am interested in understanding how efficiency bottlenecks change across workloads and architectures, and in designing methods that are tailored to those changing bottlenecks rather than assuming a single inference regime.

Broader Vision

Looking forward, I am particularly interested in efficiency challenges arising from new workloads, model architectures, and modalities. This includes agentic and long-horizon inference, multimodal models, sparse and linear attention, mixture-of-experts models, and other architectures whose computational behavior differs substantially from conventional dense autoregressive models. I believe these shifts will create new opportunities for algorithmic acceleration, but will also require a more careful understanding of workload structure, hardware utilization, memory behavior, and system-level constraints. My broader goal is to develop principled and practically effective algorithms that make increasingly capable models faster, more resource-efficient, and more economical to train and deploy. While my current focus is on inference, my earlier work in optimization and memory-efficient learning also provides a broader foundation for studying efficiency across the full lifecycle of modern foundation models.

Publications & Research Works (Last Updated: Oct. 2026)

Education

  • Ph.D. in Graduate School of AI, KAIST
    Sep. 2022 - Present
    • Advised by Prof. Eunho Yang
  • M.S. in Graduate School of AI, KAIST
    Sep. 2020 - Aug. 2022
    • Advised by Prof. Eunho Yang
  • B.S. in Computer Science and Mathematical Sciences, KAIST
    Mar. 2015 - Aug. 2020

Research Experience

  • Research Intern, Computer Architecture and Systems Lab, KAIST, Daejeon, Aug. 2019 - Dec. 2019
    • Advisor : Prof. Jaehyuk Huh
    • Low-level security techniques of Intel SGX and secure container with KVSSD
  • Research Intern, Collaborative Distributed Systems and Networking Lab, KAIST, Daejeon, Jan. 2018 - Oct. 2018
    • Advisor : Prof. Dongman Lee
    • Signal data processing for IoT task recognition and framework for task segmentation

Selected Research Projects & Collaborations

  • Efficient Foundation Models on Intel Systems
    Intel Corporation & NAVER, Sep.2024-Aug.2025

  • A Study on Conversational Large Language Models for Virtual Physicians in Patient Intake
    AITRICS, Apr.2024-May.2024

  • A Study on Optimization and Network Interpretation Method for Large-Scale Machine Learning
    National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT), Mar.2023-Feb.2027

  • A Study on Statistically and Computationally Efficient Parameter Structures for Machine Learning Algorithms
    NRF grant funded by the Korea government (MSIT), Mar.2021-Dec.2022