Computer Vision / Machine Learning
Understanding animal behavior through vision.
Research-oriented builder based in Seoul.
I study how machines can understand animal behavior from raw video, with a focus on pose estimation, action recognition, and efficient annotation.
Animal behavior analysisPose estimationVideo understandingAnnotation automation
How I Work
From reading to reproducible output
- 01Build small, reproducible machine learning and computer vision experiments.
- 02Track useful papers, datasets, tools, and benchmarks.
- 03Turn research reading into lab notes, project ideas, and working artifacts.
Selected Work
All projects Featured projects
OULAD Early At-Risk Student Prediction
archivedVirtual Learning Environment (VLE) 에서 수집하는 학생 데이터를 토대로 학습 어려움이 있는 학생을 조기 탐지
ml Open
ArchiTag – 대통령기록관 사진 AI 자동 태깅 시스템
work-in-progressVLM을 활용해 대통령 기록물의 내용 기반 메타데이터를 자동 생성하고, 이를 통해 아카이브 검색성과 운영 효율을 동시에 개선하는 연구
product Open
(Be MOre Duck) Vision Language Embodied Agent
work-in-progress라즈베리파이 기반 동물 행동 관찰 반응형 에이전트
product Open
HybridRAG – 벡터 + 그래프 기반 복합 질의 해결 RAG 시스템
archived벡터 기반 검색과 그래프 기반 검색을 결합한 하이브리드 검색 구조가, 단일 검색 방식보다 더 정확하고 관련도 높은 응답을 생성할 수 있는가
ml Open
Recent Writing
Full archive Latest lab notes
Sep 09, 2026 [Reading Note] Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureI-JEPA는 pixel reconstruction이나 hand-crafted augmentation invariance에 의존하지 않고, 보이는 context로부터 가려진 영역의 abst... Sep 01, 2026 [Reading Note] Learning Transferable Visual Models From Natural Language Supervision Aug 18, 2026 [Reading Note] VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingVideoMAE는 plain ViT에 high-ratio tube masking을 적용해 비디오 reconstruction task를 어렵게 만들고, 작은 데이터셋에서도 강한 video r... Aug 17, 2026 [Reading Note] Masked Autoencoders Are Scalable Vision LearnersHigh-ratio random masking과 asymmetric encoder-decoder를 통해 대규모 ViT를 효율적으로 self-supervised pre-training하고,... Aug 08, 2026 [Reading Note] Is Space-Time Attention All You Need for Video Understanding?TimeSformer는 ViT를 video domain으로 확장하여 spatial·temporal attention을 분리한 Divided Space-Time Attention으로 효율적인...