Publications
2026
- EMNLPSemBridge: Language Transfer in Sparse Encoders via Multilingual Semantic BridgesProceedings of the 2026 Conference on Empirical Methods in Natural Language Processing, Oct 2026
- EMNLPSHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Information RetrievalFindings of the Association for Computational Linguistics: EMNLP 2026, Oct 2026
- EMNLPDART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning ModelsFindings of the Association for Computational Linguistics: EMNLP 2026, Oct 2026
- KBSSERA: Self-referential assessment framework for bidirectional generative commonsense reasoningKnowledge-Based Systems, 2026
- ACLTowards Privacy-Preserving Large Language Model: Text-free Inference Through Alignment and AdaptationProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Jul 2026
- ACL FindingsNeedleChain: Measuring Intact Context Comprehension Capability of Large Language ModelsFindings of the Association for Computational Linguistics: ACL 2026, Jul 2026
- SIGIRBeyond Hard Negatives: The Importance of Score Distribution in Knowledge DistillationProceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2026
- ICLRImproving Semantic Proximity in Information Retrieval through Cross-Lingual AlignmentThe Fourteenth International Conference on Learning Representations, 2026
2025
- EMNLPMetric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language ModelsProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Nov 2025
- EMNLP FindingsLimaCost: Data Valuation for Instruction Tuning of Large Language ModelsFindings of the Association for Computational Linguistics: EMNLP 2025, Nov 2025
- EMNLPThe Impact of Negated Text on Hallucination with Large Language ModelsProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Nov 2025
- ACLCall for Rigor in Reporting Quality of Instruction Tuning DataProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Jul 2025
- ACLCross-Lingual Optimization for Language Transfer in Large Language ModelsProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Jul 2025
- ACL FindingsSemantic Aware Linear Transfer by Recycling Pre-trained Language Models for Cross-lingual TransferFindings of the Association for Computational Linguistics: ACL 2025, Jul 2025
- NAACL FindingsFLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language ModelsFindings of the Association for Computational Linguistics: NAACL 2025, Apr 2025
- NAACL FindingsMIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation EvaluationFindings of the Association for Computational Linguistics: NAACL 2025, Apr 2025
- NAACL FindingsFind the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language ModelsFindings of the Association for Computational Linguistics: NAACL 2025, Apr 2025
- COLINGMIGRATE: Cross-Lingual Adaptation of Domain-Specific LLMs through Code-Switching and Embedding TransferProceedings of the 31st International Conference on Computational Linguistics, Jan 2025
2024
- LREC-COLINGLeveraging Pre-existing Resources for Data-Efficient Counter-Narrative Generation in KoreanProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), May 2024
- LREC-COLINGDetecting Critical Errors Considering Cross-Cultural Factors in English-Korean TranslationProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), May 2024
- EACL FindingsGenerative Interpretation: Toward Human-Like Evaluation for Educational Question-Answer Pair GenerationFindings of the Association for Computational Linguistics: EACL 2024, Mar 2024
- EACL FindingsHyper-BTS Dataset: Scalability and Enhanced Analysis of Back TranScription (BTS) for ASR Post-ProcessingFindings of the Association for Computational Linguistics: EACL 2024, Mar 2024
- IEEEExploiting hanja-based resources in processing korean historic documents written by common literatiIEEE Access, 2024
2023
- EMNLPKEBAP: Korean Error Explainable Benchmark Dataset for ASR and Post-processingProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023
- EMNLPPost-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded ConversationsProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023
- EMNLPCHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme IngredientsThe 2023 Conference on Empirical Methods in Natural Language Processing, 2023
- IJCNLP-AACLInformative Evidence-guided Prompt-based Fine-tuning for English-Korean Critical Error DetectionProceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 2023
- IWSLTImproving formality-sensitive machine translation using data-centric approaches and prompt engineeringProceedings of the 20th International Conference on Spoken Language Translation (IWSLT 2023), 2023
- ACL DemoPEEP-Talk: A Situational Dialogue-based Chatbot for English EducationProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), 2023
- ESWA
- IEEEUncovering the Risks and Drawbacks Associated with the Use of Synthetic Data for Grammatical Error CorrectionIEEE Access, 2023
- Mathematics
2022
- COLINGQUAK: A Synthetic Quality Estimation Dataset for Korean-English Neural Machine TranslationProceedings of the 29th International Conference on Computational Linguistics, 2022
- NAACL FindingsA dog is passing over the jet? a text-generation dataset for korean commonsense reasoning and evaluationFindings of the Association for Computational Linguistics: NAACL 2022, 2022
- LRECEmpirical Analysis of Noising Scheme based Synthetic Data Generation for Automatic Post-editingProceedings of the Thirteenth Language Resources and Evaluation Conference, 2022
- LRECPriming Ancient Korean Neural Machine TranslationProceedings of the Thirteenth Language Resources and Evaluation Conference, 2022
- KBSPU-GEN: Enhancing generative commonsense reasoning for language models with human-centered knowledgeKnowledge-Based Systems, 2022
- IEEEK-nct: Korean neural grammatical error correction gold-standard test set using novel error type classification criteriaIEEE Access, 2022
- IEEEPlain Template Insertion: Korean-Prompt-Based Engineering for Few-Shot LearnersIEEE Access, 2022
- Appl. Sci.BERTOEIC: Solving TOEIC Problems Using Simple and Efficient Data Augmentation Techniques with Pretrained Transformer EncodersApplied Sciences, 2022
- Appl. Sci.
- IEEEAI for Patents: A Novel Yet Effective and Efficient Framework for Patent AnalysisIEEE Access, 2022
- MathematicsReturn on Advertising Spend Prediction with Task Decomposition-Based LSTM ModelMathematics, 2022
- IEEE
- MathematicsDense-to-question and sparse-to-answer: Hybrid retriever system for industrial frequently asked questionsMathematics, 2022
- IEEEMimicking Infants’ Bilingual Language Acquisition for Domain Specialized Neural Machine TranslationIEEE Access, 2022
- IEEE
2021
- NAACL IndustryShould we find another model?: Improving neural machine translation performance with ONE-piece tokenization method without model modificationProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Papers, 2021
- IEEE
- Appl. Sci.Comparative analysis of current approaches to quality estimation for neural machine translationApplied Sciences, 2021