Natural & Nature Language Processing Lab
@ HUFS
We are the team focusing on Artificial Intelligence for Natural & Nature Language Processing.
We are looking for self-motivated graduate and undergraduate students to join our lab.
If you are interested in our lab, please don't hesitate to reach out to us!
우리 연구실은 자연어처리 및 생명정보학을 위한 인공지능을 연구하고있습니다.
석사/박사 대학원생, 학부연구생 지원을 언제나 환영합니다.
여기의 이메일로 연락주시기 바랍니다.
Natural Language Processing (NLP) is a field of artificial intelligence that enables computers to understand, interpret, and learn patterns from human language. By identifying meaningful structures in human language NLP models can generate language, predict information, and analyze linguistic data. These capabilities support a wide range of applications, including translation, question answering, text generation, sentiment analysis, and information extraction.
Our research focuses on adapting and specializing NLP models for specific domains, including scientific literature, clinical records, and culture/domain-specific language models. Our overarching goal is that supporting people to understand each other regardless of their domain or cultural differences via lens of technology.
자연어처리(Natural Language Processing, NLP)는 컴퓨터가 인간의 언어를 이해하고 해석하며, 그 안에 존재하는 패턴을 학습할 수 있도록 하는 인공지능 분야입니다. 자연어처리 모델은 인간의 언어에서 의미 있는 구조를 찾아내고, 이를 바탕으로 언어를 생성하거나 정보를 예측하고 언어 데이터를 분석합니다. 이러한 기술은 번역, 질의응답, 텍스트 생성, 감성 분석, 정보 추출 등 다양한 분야에 활용됩니다.
우리 연구실은 과학 문헌, 임상 기록, 특정 문화적 맥락이나 전문 도메인 특성ㅇㄹ 반영한 언어모델 개발과 같은 특정 분야에 자연어처리 모델을 적용하고 특화하는 연구를 수행합니다. 궁극적으로 분야 또는 문화적 배경에 상관없이 서로를 이해하는 것을 지원하는 기술을 개발하고자 합니다.
Bioinformatics is an interdisciplinary field that uses computational methods to understand and analyze the language of nature encoded in DNA, RNA, and protein sequences. By learning meaningful patterns and relationships within biological sequences, computational models can reveal how genetic information is organized and how it contributes to biological functions and processes. These models support a wide range of tasks, including predicting the functions of genes/proteins, identifying genomic signatures (DNA fingerprinting), and predicting protein–protein interactions.
Our research focuses on developing and applying machine learning and LLM-based approaches to biological sequence analysis. We aim to build models that can learn complex biological patterns and use them to generate meaningful predictions and insights into genes, proteins, and their interactions. Our overarching goal is to contribute to human health by understanding and harnessing the language of life.
생명정보학(Bioinformatics)은 DNA, RNA, 단백질 서열에 기록된 자연의 언어를 계산적 방법으로 이해하고 분석하는 융합 학문 분야입니다. 계산 모델은 생물학적 서열에 존재하는 의미 있는 패턴과 관계를 학습함으로써 유전 정보가 어떻게 구성되어 있으며, 생물학적 기능과 과정에 어떻게 기여하는지를 밝혀낼 수 있습니다. 이러한 모델은 유전자와 단백질의 기능 예측, 유전체 특징 또는 DNA 지문 분석, 단백질–단백질 상호작용 예측 등 다양한 과제에 활용됩니다.
우리 연구실은 생물학적 서열 분석을 위한 머신러닝 및 대규모 언어 모델(LLM) 기반 방법론을 개발하고 적용하는 연구를 수행합니다. 복잡한 생물학적 패턴을 학습하고, 이를 바탕으로 유전자와 단백질, 그리고 이들의 상호작용에 관한 의미 있는 예측과 통찰을 제공할 수 있는 모델을 구축하는 것을 목표로 합니다. 궁극적으로는 생명의 언어를 이해하고 활용함으로써 인류의 건강 증진에 기여하고자 합니다.
AI: Artificial Intelligence, ML: Machine Learning , NLP: Natural Language Processing, LLM: Large Language Model
BI: Bioinformatics, MG: Metagenomics, HM: Human Microbiome, EHR: Electronic Health Record
TBD