I am an AI Researcher at AITRICS, where I fine-tune and deploy ASR in production under domain shift, improve the LLM summarization pipeline built on its transcripts, and compress LLMs to cut serving costs. I am broadly interested in making multimodal systems efficient and reliable in deployment, from model compression and pruning to evaluating how they behave across diverse users, languages, and real-world conditions. To address the inefficiencies I experienced during research, I developed and open-sourced Evalmetry, a framework that collects research data during model evaluation.
Before AITRICS, I received my M.S. in Artificial Intelligence from POSTECH (August 2025), advised by Prof. Jaeho Lee at the Efficient Learning Lab, where I worked on language-guided multimodal learning, including text-guided image compression. I previously worked with Prof. Tae-Hyun Oh on audio captioning. I earned my B.S. in Electrical and Electronics Engineering from Chung-Ang University.
I work on multimodal learning, with a focus on building vision–language and speech systems that are both efficient and reliable.
(* means ‘equal contribution’)
Minkyu Kim*, Vincent-Daniel Yun*, Youngrae Kim*, Suin Cho, Woosang Lim, Sunwoo Lee, “Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning”, 2026
Minkyu Kim, Juhwan Choi, YoungBin Kim, “Not Safe for All: Auditing the Dialect Penalty in Text-to-Image Safety Pipelines”, EMNLP (Findings), 2026
Vincent-Daniel Yun*, Youngrae Kim*, Woosang Lim*, Youngjin Heo, Minkyu Kim, Sunwoo Lee, “Locality-Aware Redundancy Pruning for LLM Depth Compression”, EMNLP (Main), 2026
Jean Seo*, Minkyu Kim*, Jeonguk Lee, Jisoo Jung, Wooseok Han, Eunho Yang, “When Multiple Scripts Matter: Evaluating ASR in Clinical Settings”, Interspeech, 2026
Hagyeong Lee*, Minkyu Kim*, Jun-Hyuk Kim, Seungeon Kim, Dokwan Oh, Jaeho Lee, “Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity”, ICML, 2024
Minkyu Kim*, Kim Sung-Bin*, Tae-Hyun Oh, “Prefix Tuning for Automated Audio Captioning”, ICASSP (Oral), 2023