Junu's Lab
LAB ONLINE Focus Computer Vision / ML / Animal Behavior
Weekly Lab Brief

Animal Behaviour Analysis Weekly Brief — 2026.08.10.

Completed
Type
weekly-brief
Date
Aug 10, 2026
Domain
Methods
Research tracking routine
Selected for review:

Animal Behaviour Analysis Weekly Research Briefing

기준일: 2026년 8월 10일
우선 조사 기간: 2026년 7월 27일~8월 10일
분야: Animal Behaviour Analysis / Computational Ethology / Animal Pose Estimation / Animal Behavior Recognition
이번 주 핵심: 이번 주에는 사용자의 관심사와 직접 맞닿는 논문이 하나 나왔습니다. Promptable Animal Pose Tracking Across Species 는 동물마다 수백 장의 keypoint annotation을 다시 만드는 대신, 한 reference frame에서 원하는 keypoint를 지정하고 나머지 영상으로 전파 하는 방향을 제시합니다. supervised route와 training-free unsupervised route를 모두 제공하며, foundation-model feature를 이용해 cross-species pose tracking을 수행합니다. 2026년 8월 5일 arXiv에 공개됐고 ECCV 2026 CV4Ecology Workshop 발표 논문입니다. ( arXiv )

1. 이번 주 한눈에 보기

이번 주 가장 중요한 변화는 animal pose labeling 자체를 줄이는 방법이 foundation model 기반 promptable tracking 형태로 구체화되었다는 점 입니다. Promptable Animal Pose Tracking은 영상 한 프레임에서 사용자가 관심 keypoint를 지정하면, DINOv3·BioCLIP·diffusion feature 같은 pretrained representation을 이용해 나머지 frame으로 keypoint를 전달합니다. 연구진은 supervised route와 별도로 task-specific training이 필요 없는 unsupervised correspondence route 도 제안했습니다. ( arXiv ) 특히 이 연구는 사용자의 labeling 자동화 관심과 직접 연결됩니다. DeepLabCut/SLEAP처럼 여러 frame을 직접 annotation해 species-specific pose estimator를 학습하는 대신, “한 프레임을 prompt로 주고 tracking으로 annotation을 확장한다” 는 방향입니다. 논문은 실제 연구자가 predefined skeleton 전체가 아니라 관심 있는 keypoint 몇 개만 지정하고 싶어 한다는 문제를 명시적으로 다룹니다. ( arXiv ) 또한 7월 28일에는 Comparison of multiple video tracking-based behavioral summary approaches for compound discrimination 의 v2가 공개됐습니다. 이 연구의 중요한 메시지는 복잡한 unsupervised behavior representation이 항상 단순 tracking summary보다 압도적으로 낫지는 않다는 점입니다. 검색 가능한 논문 설명에서는 treatment/dosage discrimination 성능이 여러 behavioral summary model 사이에서 유사했다고 보고합니다. ( bioRxiv ) 도구 측면에서는 SLEAP이 현재 PyTorch backend와 sleap-io , sleap-nn 으로 모듈화되어 있고, GUI 수준에서 human-in-the-loop labeling workflow를 제공하고 있습니다. 최신 신작은 아니지만 이번 주 promptable tracking 흐름과 비교할 기준점으로 적절합니다. ( TALMO Lab )

이번 주 labeling 자동화 관점의 핵심

기존 흐름은 대체로 다음과 같습니다.
많은 frame 수동 labeling

Animal-specific pose model training

전체 video inference

오류 frame 수정
이번 주 핵심 연구가 제안하는 흐름은 더 가볍습니다.
Reference frame 1장

사용자가 원하는 keypoint prompt

Foundation-model dense feature

Cross-frame correspondence

전체 video keypoint trajectory
즉, pose estimation 문제 일부를 “supervised keypoint detection”에서 “visual correspondence tracking” 문제로 재정의 합니다. ( arXiv )

이번 주 반드시 볼 자료

Promptable Animal Pose Tracking Across Species 이 논문은 지금까지 추적해온 주제 중에서도 다음 항목과 동시에 연결됩니다.
  • labeling 자동화
  • scarce annotation
  • cross-species generalization
  • keypoint sequence
  • animal pose tracking
  • vision foundation model
  • wildlife video
  • downstream behavior recognition
따라서 이번 주에는 두 논문을 동일한 비중으로 읽기보다 Promptable Animal Pose Tracking을 주 논문으로 깊게 읽고 , behavioral-summary comparison은 결과와 Discussion 위주로 보는 것이 적절합니다.

2. 이번 주 핵심 트렌드

2.1 “Few-label pose estimation”에서 “Promptable pose tracking”으로

무엇이 바뀌었는가

기존 scarce-label 방법은 적은 annotation으로 pose estimator를 잘 fine-tuning하는 문제로 접근했습니다. Promptable Animal Pose Tracking은 다른 문제 설정을 사용합니다.
Video
+
Reference frame의 K개 keypoint

Foundation model

Reference ↔ target feature correspondence

K개 keypoint를 모든 frame으로 propagate
즉, 새 species마다 pose detector를 처음부터 학습해야 한다는 가정을 완화 합니다. 논문은 단 하나의 annotated reference frame을 입력으로 사용해 나머지 video frame의 keypoint trajectory를 추정합니다. ( arXiv )

왜 중요한가

animal pose에서 가장 큰 병목 중 하나는 morphology가 species마다 달라 standard skeleton을 만들기 어렵다는 점입니다. 논문 역시 cross-species appearance와 morphology 차이, annotation 비용을 핵심 문제로 지적합니다. ( arXiv )

Animal Behaviour Analysis 연결

행동 분석 전체 pipeline을 다음처럼 만들 수 있습니다.
1 frame manual keypoint

Promptable tracking

Full-video keypoint sequence

movement / trajectory feature

behavior segmentation

behavior classifier / forecasting
사용자의 장기 연구 흐름과 상당히 잘 맞습니다.

내가 지금 신경 써야 하는가

High

2.2 Supervised와 unsupervised를 경쟁 관계가 아니라 역할 분담으로 사용

Promptable APT에는 두 route가 있습니다.

Supervised route

Reference image + target image

Frozen foundation-model features

Keypoint prompt encoder

Feature projection

Correspondence matcher

Target-frame keypoint
foundation backbone 전체를 학습하는 것이 아니라 keypoint prompt encoder와 feature projection, matcher 정도를 학습합니다. 논문에서는 foundation backbone을 frozen 상태로 두고 이 구성요소들을 50 epoch 학습했다고 설명합니다. ( arXiv ) 장점은 occlusion, ambiguity, 큰 appearance 변화에서 정확도가 좋다는 것 입니다. ( arXiv )

Unsupervised route

Reference image + target image

Frozen foundation-model dense feature

Pixel-level correspondence

Bounding-box constraint

Drift correction

Target keypoint
task-specific pose training이 필요 없습니다. 장점은 unseen species와 category shift에 상대적으로 강한 cross-species generalization 입니다. 논문의 leave-one-family-out 실험에서는 supervised model이 가까운 category가 없는 경우 negative transfer를 보이는 반면 unsupervised route가 더 안정적인 경우도 관찰됐습니다. ( arXiv )

중요한 연구적 포인트

이는 흥미로운 trade-off입니다.
Supervised = accuracy 중심Unsupervised = generalization 중심
완전히 한쪽이 다른 쪽을 대체하는 구조가 아닙니다. ( arXiv )

내가 지금 신경 써야 하는가

High 특히 향후 연구에서 다음 비교가 가능합니다.
DLC/SLEAP supervised
vs
Promptable supervised
vs
Promptable unsupervised
그리고 비교 지표를 annotation frame 수와 tracking accuracy로 잡을 수 있습니다.

2.3 Vision Foundation Model이 animal pose의 feature extractor로 들어오기 시작

논문은 dense per-pixel feature extractor 후보로 다음 계열을 다룹니다.
  • DINOv3
  • BioCLIP
  • CleanDIFT
  • Diffusion Hyperfeatures
즉 foundation model을 단순 image classifier encoder로 사용하는 것이 아니라 keypoint correspondence를 위한 dense representation 으로 활용합니다. ( arXiv ) 이는 중요한 변화입니다. 기존에는:
Foundation model
→ image/video embedding
→ behavior classification
을 많이 생각했다면, 이제는:
Foundation model
→ spatial correspondence
→ keypoint tracking
→ behavior sequence
로도 연결할 수 있습니다.

Animal Behaviour Analysis 연결

사용자 연구에서는 CLIP 평균 embedding을 바로 behavior classifier에 넣는 것뿐 아니라 foundation feature를 tracking과 annotation automation 단계 에서 먼저 활용하는 방향도 고려할 수 있습니다.

내가 지금 신경 써야 하는가

High

2.4 단순한 behavior summary를 무시하면 안 됨

7월 28일 v2가 공개된 behavioral summary 비교 연구는 video tracking에서 얻은 정보를 여러 방식으로 요약하여 compound/treatment discrimination에 활용합니다. bioRxiv 검색 결과에서 저자들은 여러 behavioral summary model 사이의 discrimination 성능이 유사한 경우를 보고했습니다. ( bioRxiv ) 따라서 다음과 같은 가정은 피하는 것이 좋습니다.
Keypoint-MoSeq
>
trajectory statistics
또는
Transformer
>
handcrafted feature
복잡도가 높다는 이유만으로 자동으로 더 좋은 representation이 되는 것은 아닙니다.

Animal Behaviour Analysis 연결

BMOD/Duck에서 다음 순서가 더 논리적입니다.
Baseline 1
centroid trajectory statistics

Baseline 2
pose statistics

Baseline 3
unsupervised pose state

Baseline 4
learned video embedding
이후 성능을 비교해야 합니다.

내가 지금 신경 써야 하는가

High

3. 추천 논문 1~2편

3.1 Promptable Animal Pose Tracking Across Species

Title: Promptable Animal Pose Tracking Across Species
Authors: Le Li, Daniela Ivanova, Nicolas Pugeault
Affiliation: University of Glasgow
공개: 2026년 8월 5일
Venue 상태: ECCV 2026 Workshop on CV4Ecology 발표 예정
논문: arXiv:2608.04995 ( arXiv )

분야

  • Animal Pose Tracking
  • Vision Foundation Model
  • Few-shot / Low-label Learning
  • Unsupervised Correspondence
  • Cross-species Generalization

핵심 문제

동물 pose tracking은 human pose에 비해 다음 문제가 큽니다.
종별 morphology 차이
+
행동 차이
+
annotation 부족
+
standard skeleton 부족
기존 animal pose model은 많은 annotation을 요구하거나, 특정 species에 지나치게 특화될 수 있습니다. 반대로 general point tracker는 임의의 point를 tracking할 수 있지만 animal anatomy prior가 부족할 수 있습니다. 논문은 이 간극을 해결하는 것을 목표로 합니다. ( arXiv )

핵심 아이디어

사용자가 video의 한 reference frame만 annotation 합니다. 예:
Duck frame 1

● beak
● head
● neck
● left wing
● right wing
● tail
그다음 foundation-model dense feature를 이용해 동일한 anatomical point를 다른 frame에서 찾습니다.
Reference keypoint

Reference feature vector

Target-frame dense feature map

Correspondence matching

Target keypoint
이 과정을 전체 video로 반복합니다. ( arXiv )

모델 구조

입력

RGB video
+
Reference frame 1개
+
Reference keypoints K개

공통 처리

Reference frame
Target frame

Foundation-model feature extractor

Dense feature map
논문은 DINOv3, BioCLIP, CleanDIFT, Diffusion Hyperfeatures 같은 pretrained representation을 검토합니다. ( arXiv )

Supervised route

Reference feature
+
keypoint prompt

Keypoint Prompt Encoder

Projected feature

Matcher

Target keypoint

Unsupervised route

Reference feature

Dense correspondence

Bounding-box restriction

Drift correction

Target keypoint

출력

T frames × K keypoints × (x,y)
즉 downstream에서 바로 keypoint sequence로 사용할 수 있습니다. ( arXiv )

실험 데이터셋

APTv2

APTv2에는 30종, 15 family에 걸친 animal video가 포함됩니다. 각 clip은 15 frame이며, 논문에서는 첫 번째 frame의 keypoint를 reference annotation으로 사용해 나머지 14 frame으로 전파합니다. ( arXiv )

TigDog

동물 pose generalization 평가에 사용하는 기존 benchmark입니다. Promptable APT는 APTv2와 TigDog 모두에서 평가됩니다. ( arXiv )

성능 또는 주요 결과

수치 하나만 떼어 SOTA라고 판단하는 것보다는 이 논문의 trade-off 결과 가 더 중요합니다.

Supervised

  • tracking accuracy가 더 높음
  • occlusion에 강함
  • appearance 변화에 더 안정적

Unsupervised

  • task-specific training 불필요
  • unseen species에 상대적으로 안정적
  • category shift에 강함
연구진은 foundation backbone을 frozen 상태로 유지하고 supervised route의 일부 모듈만 학습했음에도 강한 tracking 성능을 보고합니다. ( arXiv )

기존 방법과의 차이

논문의 appendix는 대표적인 animal tracking 방법과 다음처럼 비교합니다.
DeepLabCut
→ supervised

ScarceNet
→ pseudo-label-based supervised

3D-MuPPET
→ supervised multi-view

Promptable APT
→ supervised / unsupervised
→ foundation model
→ 30 species
특히 fixed predefined skeleton보다 user-selected keypoint 라는 것이 중요한 차이입니다. ( arXiv )

한계점

첫째, APTv2 clip이 15 frame으로 상당히 짧습니다. 논문에서도 APTv2 frame은 일정 FPS로 연속 sampling된 것이 아니라 동작 변화가 보이도록 수동 선택되었다고 설명합니다. 따라서 몇 시간짜리 continuous video에서 long-term drift가 어떻게 누적되는지는 별도 검증이 필요 합니다. ( arXiv ) 둘째, bounding box가 필요합니다. 셋째, supervised route는 cross-species generalization에서 training distribution에 따라 negative transfer가 발생할 수 있습니다. ( arXiv ) 넷째, 현재 논문에서는 pose tracking이 목표이며 행동 분류까지 end-to-end로 연결하지는 않습니다.

내 관심사와의 연결

매우 높음 특히 다음 연구 질문으로 연결할 수 있습니다.
“한 frame keypoint prompt만으로 장시간 single-animal video의 pose sequence를 생성하고, 이를 behavior embedding/classification에 사용할 수 있는가?”
이 주제는 충분히 narrow하면서도 다음을 묶습니다.
  • automated annotation
  • foundation model
  • animal pose
  • long video
  • behavior recognition

읽기 우선순위

High — 이번 주 1순위

읽을 때 집중할 부분

  1. Reference frame 선택이 성능에 미치는 영향
  2. Foundation feature별 성능
  3. supervised vs unsupervised 차이
  4. drift correction
  5. leave-one-family-out experiment
  6. bounding-box dependence
  7. temporal robustness
  8. annotation 절감량을 어떤 방식으로 평가했는지
특히 appendix의 Temporal Robustness Leave-one-out analysis 를 놓치지 않는 것이 좋습니다.

구현 난이도

전체 재현

어려움 Foundation-model feature extraction과 dense correspondence 구현이 필요합니다.

축소 구현

보통 이미 존재하는 point tracker 또는 DINO feature correspondence를 사용하면 핵심 아이디어만 실험할 수 있습니다.

현재 수준에서 필요한 선행 학습

  • PyTorch tensor indexing
  • ViT feature map
  • cosine similarity
  • feature correspondence
  • bounding box coordinate
  • bilinear interpolation
  • nearest-neighbor matching
  • video frame extraction
Transformer를 처음부터 구현할 필요는 없습니다.

포트폴리오 포인트

Vision foundation model의 dense representation을 이용해 한 reference frame의 animal keypoint annotation을 전체 video로 전파하고, manual labeling 비용과 tracking drift의 trade-off를 분석했습니다.

3.2 Comparison of Multiple Video Tracking-Based Behavioral Summary Approaches for Compound Discrimination

공개: v2, 2026년 7월 28일
상태: bioRxiv preprint
논문: bioRxiv 2026.07.20.739643v2 ( bioRxiv )

분야

  • behavior representation
  • animal tracking
  • trajectory analysis
  • unsupervised behavior segmentation

핵심 문제

동물 tracking을 완료했다고 해서 행동 분석이 끝나는 것은 아닙니다. 예를 들어 tracking sequence가 있을 때 다음 중 무엇을 사용할 수 있습니다.
speed mean
distance traveled
zone occupancy
turning

vs

unsupervised state

vs

behavioral syllable
어느 representation이 실험 조건의 차이를 가장 잘 나타내는가가 핵심 문제입니다.

핵심 아이디어

여러 behavioral summary method를 동일 downstream discrimination 문제에 적용해 비교합니다. 검색 가능한 논문 설명에서는 treatment/dosage discrimination 성능이 여러 모델 사이에서 비슷한 경향을 보였다고 보고합니다. 즉 더 복잡한 representation이 항상 압도적 이득을 주지는 않았습니다. ( bioRxiv )

내 관심사와의 연결

High 현재 Duck 연구에서도 반드시 고려해야 하는 문제입니다. 다음 순서로 baseline을 만들 수 있습니다.
1. Trajectory statistics
2. Pose statistics
3. Keypoint-MoSeq
4. CLIP/video embedding
5. Fusion
이렇게 해야 모델 성능 향상의 원인을 설명할 수 있습니다.

읽기 우선순위

Medium~High 이번 주에는 Methods 전체보다 다음을 먼저 보는 것이 좋습니다.
  • 비교한 representation 종류
  • downstream evaluation 방법
  • complexity 대비 성능
  • Discussion

구현 난이도

  • trajectory feature: 쉬움
  • Keypoint-MoSeq: 보통
  • 전체 paper reproduction: 어려움

4. 추천 GitHub repo / tool

이번 주에는 억지로 최신 repo 3개를 채우지 않고 실제로 지금 쓸 가치가 있는 2개만 권합니다.

4.1 SLEAP 1.5+

공식 사이트: SLEAP SLEAP은 현재 GUI의 human-in-the-loop annotation과 single/multi-animal pose estimation을 지원하며, backend가 PyTorch 중심으로 전환되었습니다. sleap-io sleap-nn 으로 데이터 처리와 neural-network workflow도 분리되어 있습니다. ( TALMO Lab )

무엇을 하는가

Video

manual label

model training

prediction

human correction

retraining

연구적으로 중요한 이유

Promptable APT와 비교할 좋은 baseline입니다. 연구 질문을 다음처럼 잡을 수 있습니다.
SLEAP
N manually annotated frames

vs

Promptable tracking
1 manually annotated frame

설치/실행 난이도

보통 이하 공식 설치법은 현재 uv 와 pip를 모두 지원합니다. ( TALMO Lab )

당장 실험해볼 수 있는 부분

Duck 영상 30초 정도로:
  • 5~10 frame labeling
  • 작은 pose model 학습
  • 전체 영상 inference
  • 오류 frame 수 측정
이후 Promptable 방식과 비교할 baseline으로 보관하면 됩니다.

4.2 APTv2

논문/데이터: APTv2 최신 dataset은 아니지만 이번 주 신작 논문의 핵심 benchmark이므로 다시 볼 가치가 있습니다. APTv2는 30 animal species, 2,749 video clip, 총 41,235 frame과 84,611 animal instance의 pose/tracking annotation을 제공합니다. ( arXiv )

연구적으로 중요한 이유

다음 두 가지를 직접 평가할 수 있습니다.
In-species performance
vs
Cross-species generalization
특히 labeling automation 연구라면 단일 species accuracy만 보고 끝내기보다 unseen species에 keypoint propagation이 가능한가 를 볼 필요가 있습니다.

설치/실행 난이도

전체 benchmark: 어려움 일부 clip visualization: 쉬움

당장 실험할 수 있는 부분

전체 모델을 학습하지 말고 APTv2에서 서로 다른 species 5개를 골라 다음만 확인하는 것이 좋습니다.
reference frame

keypoint

다른 frame visual correspondence

5. 이번 주 미니프로젝트 제안

프로젝트 제목

One-Frame Animal Pose Propagation Baseline

목표

동물 영상에서 한 frame만 keypoint annotation 하고, foundation-model visual feature 또는 point tracker를 사용해 나머지 frame으로 keypoint를 자동 propagation합니다. 그리고 자동 생성 label과 수동 ground truth 사이의 error를 측정합니다.

배경

이번 주 Promptable Animal Pose Tracking의 핵심 연구 질문을 그대로 축소합니다. 하지만 논문 전체의 supervised model을 구현하지 않습니다. 대신 다음 질문 하나만 검증합니다.
한 장의 annotation만으로 짧은 single-animal video의 pose label을 어느 정도까지 확장할 수 있는가?

사용할 데이터

1순위

단일 Duck 영상 30~60초 정도면 충분합니다. 우선 keypoint는 3~5개만 사용합니다. 예:
beak
head
body_center
tail
foot

대체 데이터

APTv2의 single-animal clip.

사용할 모델/tool

난이도 순으로 선택합니다.

Option A — 쉬움

기존 point tracker
CoTracker
또는
TAPIR

Option B — 보통

DINO 계열 feature matching
Reference patch feature

Target feature map

cosine similarity

argmax

Option C — 어려움

Promptable APT 논문의 unsupervised route 일부 재현 현재는 A 또는 B를 권합니다.

최소 구현 범위

Step 1

30초 동물 영상 준비.

Step 2

첫 frame에서 3개 keypoint 수동 지정.
head
body
tail

Step 3

10 frame마다 ground-truth를 직접 labeling. 예:
frame 0
frame 10
frame 20
...
evaluation 용으로만 사용합니다.

Step 4

첫 frame annotation을 자동 propagation.

Step 5

각 frame에서 Euclidean distance error 계산.

Step 6

frame index에 따른 error plot 생성. 예상되는 형태:
Error
  |
  |           /
  |        /
  |     /
  |__/
  +-------------- Time
drift가 보이는지를 확인합니다.

Step 7

다음 조건 비교.
A. 첫 frame만 reference

B. 영상 중간 frame reference

C. 5초마다 reference reset

핵심 평가 지표

Pixel Error

|| prediction - GT ||

Normalized Error

animal bounding-box 크기로 normalization.

Label Saving Ratio

예를 들어:
기존: 100 frames labeling
제안: 5 frames labeling

→ manual annotation 95% 감소
단, annotation 감소율과 tracking error를 반드시 함께 보고해야 합니다.

확장 구현 범위

확장 1 — Confidence-based re-annotation

confidence가 낮은 frame만 사람이 다시 labeling합니다.
자동 tracking

low confidence

human correction

새 reference

tracking 재시작
이 단계부터 active labeling 과 직접 연결됩니다.

확장 2 — Behavior-aware sampling

행동 전환이 큰 구간만 추가 annotation합니다. 예:
Resting

Exploring
전환점에서 tracking error가 커지는지 확인합니다.

확장 3 — Downstream behavior classifier

자동 생성된 pose와 수동 pose를 각각 behavior classifier에 넣습니다. 최종적으로:
Pose error

Behavior classification error
관계를 볼 수 있습니다. 이것이 연구적으로 가장 재미있는 확장입니다.

예상 산출물

  • Notion 논문 정리
  • Jupyter notebook
  • keypoint propagation demo
  • drift plot
  • annotation saving vs error graph
  • GitHub README
  • Velog 글
  • GIF/video visualization

포트폴리오 한 줄

한 프레임의 animal keypoint annotation을 foundation-model 기반 visual tracking으로 전체 영상에 전파하고, annotation 절감률과 temporal tracking drift 간 trade-off를 정량적으로 분석했습니다.

예상 소요 시간

6~10시간 전체 Promptable APT reproduction보다 훨씬 현실적입니다.

실패 가능성이 높은 부분

가장 가능성이 높은 실패는 occlusion 이후 drift 입니다. 예:
head

wing에 가림

tracker가 wing을 head로 추적

이후 모든 frame에서 drift
두 번째 문제는 deformation입니다. 동물 몸은 rigid object가 아니기 때문에 point appearance가 크게 변할 수 있습니다.

실패했을 때 대체 목표

tracking 성능을 높이는 데 집착하지 말고 다음 실험으로 바꿉니다.
“Manual reset 간격이 pose tracking error에 미치는 영향”
예:
Reference every 1 sec
Reference every 3 sec
Reference every 5 sec
Reference only once
이것만으로도 annotation 비용과 정확도의 trade-off를 분석하는 좋은 결과가 됩니다.

6. 이번 주 읽기 루틴

1회독 — 30~40분

Promptable Animal Pose Tracking에서 다음만 봅니다.
  1. Abstract
  2. Figure 1
  3. Introduction
  4. supervised / unsupervised route 차이
  5. 전체 result table
  6. Conclusion
목표는 다음 문장을 설명할 수 있게 되는 것입니다.
“왜 animal pose estimation을 correspondence problem으로 바꾸려 했는가?”

2회독 — 1~1.5시간

이번에는 모델 구조를 봅니다. 특히 다음 pipeline을 직접 그려보십시오.
Reference frame
+
Reference keypoints

Foundation model

Dense reference feature
───────────────────────

Target frame

Foundation model

Dense target feature



Correspondence matching



Target keypoint
그리고 다음 네 가지를 확인합니다.
  • DINOv3
  • BioCLIP
  • CleanDIFT
  • Diffusion Hyperfeatures
각각 어떤 feature를 사용하는지 비교하십시오. ( arXiv )

3회독 — 내 연구 연결

이번에는 논문을 보는 것이 아니라 다음 pipeline을 설계합니다.
Camera

Motion event

Animal detection

1-frame pose prompt

Pose propagation

Pose sequence

Behavior embedding

Behavior classification
여기서 자신의 연구 질문을 하나 선택하십시오. 가장 권하는 것은:
Promptable pose tracking으로 만든 pseudo-keypoint가 downstream behavior recognition에 충분한가?
이 질문은 꽤 narrow하면서도 연구로 확장할 여지가 있습니다.

정리할 질문 3개

Q1.

왜 supervised route는 unseen family에서 unsupervised route보다 오히려 나빠질 수 있는가? → negative transfer와 learned category prior 관점에서 생각해보면 좋습니다. 논문의 leave-one-out 분석에서도 이 현상이 관찰됩니다. ( arXiv )

Q2.

Pose tracking의 error가 몇 pixel 정도까지 커져도 downstream behavior recognition에는 문제가 없을까? 이 질문은 논문이 직접 답하지 않습니다. 사용자 연구로 이어질 수 있는 좋은 open question입니다.

Q3.

Reference frame을 자동으로 선택할 수 있을까? 예:
highest confidence
+
minimal occlusion
+
representative pose
가 되는 frame을 자동으로 선택하면 annotation automation을 한 단계 더 진행할 수 있습니다.

7. 다음 주 Watchlist

추적 키워드 5개

  1. promptable animal pose tracking
  2. foundation model animal keypoint tracking
  3. active learning animal pose annotation
  4. pseudo-label animal keypoint video
  5. pose tracking downstream behavior recognition

추적할 repo/tool 3개

1. SLEAP

특히 PyTorch-based sleap-nn 과 human-in-the-loop workflow 변화를 추적합니다. ( TALMO Lab )

2. APTv2

Promptable APT 관련 code/checkpoint가 공개되면 benchmark reproduction에 활용합니다.

3. Promptable Animal Pose Tracking

현재 arXiv 본문에서는 공식 GitHub code repository가 명확히 연결되어 있지 않습니다. 따라서 공식 code 공개 여부를 다음 주에도 확인할 가치가 큽니다. ( arXiv )

다음 주 open question 3개

  1. Promptable Animal Pose Tracking의 공식 code와 pretrained weight가 공개되는가?
  2. 15-frame APTv2가 아니라 수백~수천 frame의 continuous animal video에서도 drift가 안정적인가?
  3. promptable pose label을 그대로 pseudo-label로 사용해 SLEAP/DeepLabCut을 재학습하면 annotation 효율을 더 높일 수 있는가?

Notion / Velog / 포트폴리오 게시용 자료

제목

동물 Pose Labeling을 한 프레임으로 줄일 수 있을까? — Promptable Animal Pose Tracking Across Species

요약문

2026년 8월 공개된 Promptable Animal Pose Tracking Across Species 는 animal pose estimation의 annotation 문제를 다른 방식으로 접근한다. 여러 frame에 keypoint를 직접 labeling해 pose model을 학습하는 대신, 하나의 reference frame에서 사용자가 원하는 keypoint를 지정하고 vision foundation model의 dense representation을 이용해 이를 나머지 video frame으로 전파한다. 연구진은 supervised route와 unsupervised route를 모두 제안한다. supervised 방식은 keypoint prompt encoder와 correspondence matcher를 학습해 높은 tracking accuracy를 얻는 데 초점을 두는 반면, unsupervised 방식은 task-specific training 없이 foundation-model feature correspondence를 사용해 unseen species에 대한 generalization을 확보한다. 이 구조는 animal behavior analysis의 labeling pipeline을 manual labeling → model training → inference 에서 single-frame prompt → automatic propagation → human correction 으로 바꿀 가능성이 있다. 이번 주 실험에서는 전체 모델을 재현하지 않고, 짧은 동물 영상에서 첫 frame의 3~5개 keypoint를 point tracker 또는 DINO feature matching으로 propagation한다. 이후 수동 ground truth와 비교해 tracking error가 시간에 따라 얼마나 누적되는지 분석한다.

태그

#AnimalBehaviourAnalysis#ComputationalEthology#AnimalPoseEstimation#AnimalPoseTracking#AutomatedAnnotation#PseudoLabeling#FoundationModel#DINO#SLEAP#ComputerVision

GitHub README 초안

# One-Frame Animal Pose Propagation

## Overview

This project explores whether animal pose annotations can be propagated across a video from a single manually annotated reference frame.

The project is inspired by *Promptable Animal Pose Tracking Across Species*, which uses dense visual representations from vision foundation models for animal keypoint correspondence.

## Research Question

How much manual animal pose annotation can be reduced through prompt-based keypoint tracking before temporal drift becomes too large?

## Pipeline

1.Load a short animal video.
2.Manually annotate several keypoints in one reference frame.
3.Extract visual features or initialize a point tracker.
4.Propagate each keypoint across subsequent frames.
5.Manually annotate sparse evaluation frames.
6.Measure keypoint tracking error.
7.Analyze error growth over time.

## Keypoints

Example:

-head
-beak
-body center
-tail

## Baselines

-point tracker
-DINO feature correspondence
-periodic manual reference reset

## Evaluation

-pixel error
-normalized keypoint error
-temporal drift
-tracking failure rate
-manual annotation reduction

## Experiments

### Experiment 1

Single reference frame.

### Experiment 2

Reference reset every 5 seconds.

### Experiment 3

Reference reset every 2 seconds.

## Extension

Use low-confidence frames as candidates for human correction.

The corrected frame becomes a new reference frame, creating a human-in-the-loop annotation pipeline.

## Long-Term Goal

Video
→ promptable pose tracking
→ pose sequence
→ behavior recognition
→ behavior forecasting

## Portfolio Statement

Implemented a one-frame animal pose propagation pipeline and quantified the trade-off between manual annotation reduction and temporal tracking drift.

간결한 블로그 초안

동물 Pose Labeling을 한 프레임으로 줄일 수 있을까?

Animal pose estimation을 사용하려면 일반적으로 여러 영상 frame에 keypoint를 직접 annotation해야 한다. 새로운 동물 종이나 촬영 환경으로 바뀔 때마다 이 과정을 반복해야 하기 때문에 pose estimation은 동물 행동 분석 pipeline에서 큰 labeling 병목이 된다. 2026년 8월 공개된 Promptable Animal Pose Tracking Across Species 는 이 문제를 pose detection이 아니라 visual correspondence 문제로 접근한다. 사용자는 video의 한 reference frame에서 원하는 keypoint만 지정한다. 예를 들어 오리 영상이라면 head, beak, body, tail 네 점을 지정할 수 있다. 이후 vision foundation model로 reference frame과 target frame의 dense visual feature를 추출한다. reference keypoint에 대응하는 feature와 가장 유사한 target-frame 위치를 찾으면 해당 keypoint를 나머지 영상으로 전파할 수 있다. 논문은 supervised 방식과 unsupervised 방식을 모두 제안한다. Supervised 방식에서는 keypoint prompt encoder와 correspondence matcher를 학습한다. 이 방식은 occlusion이나 큰 appearance 변화에서 높은 정확도를 목표로 한다. 반면 unsupervised 방식에서는 별도의 animal pose training 없이 pretrained foundation-model representation 자체를 사용한다. 정확도에서는 supervised 방식에 뒤질 수 있지만 새로운 species에 대한 generalization이 더 안정적인 경우가 관찰되었다. 이 결과에서 중요한 점은 두 방법이 완전히 경쟁하는 것이 아니라 서로 다른 목적을 가진다는 것이다.
Supervised
→ Accuracy

Unsupervised
→ Cross-species generalization
이 연구는 labeling automation이라는 관점에서도 흥미롭다. 기존 workflow가
많은 frame 수동 labeling
→ pose model training
→ video inference
였다면, promptable tracking은
한 frame labeling
→ automatic propagation
→ 필요한 frame만 human correction
이라는 구조를 만들 수 있기 때문이다. 이번 주에는 논문 전체 모델을 재현하기보다 이 핵심 아이디어만 실험한다. 짧은 animal video에서 첫 frame의 3~5개 keypoint를 지정하고 point tracker 또는 foundation-model feature matching으로 나머지 frame에 전파한다. 이후 일정 간격으로 만든 수동 ground truth와 비교해 tracking error가 시간에 따라 얼마나 증가하는지를 측정한다. 다음 단계에서는 error가 커지거나 confidence가 낮은 frame만 사람이 수정하도록 만들 수 있다. 이렇게 되면 단순 pose tracking experiment가 아니라,
automatic pose labeling
→ uncertainty detection
→ human review
→ pseudo-label generation
으로 이어지는 human-in-the-loop animal annotation pipeline이 된다.

이번 주 최종 우선순위

지금 바로 할 것: Promptable Animal Pose Tracking 논문의 Figure 1 → Method → unsupervised route → cross-species/temporal robustness 순서로 읽고, one-frame propagation baseline을 설계하는 것을 권합니다. 이 주제는 지난주의 behavior forecasting보다 현재의 labeling 자동화 관심과 더 직접적으로 맞닿아 있습니다. ( arXiv ) 나중에 봐도 되는 것: behavioral-summary comparison은 중요한 논문이지만, 이번 주에는 tracking representation에 대한 비교 결과와 Discussion만 확인해도 충분합니다. ( bioRxiv ) 이번 주 신뢰도 구분: Promptable APT는 ECCV 2026 workshop 발표 예정이지만 현재 공개본은 arXiv v1이고, behavioral-summary comparison은 bioRxiv preprint입니다. 따라서 둘 다 최종적으로 확정된 범용 SOTA로 받아들이기보다는 새 연구 방향과 실험 설계 근거 로 사용하는 것이 적절합니다. ( arXiv )

[Paper Quest] 주간 논문 분석 및 최종 선택 리포트

분석 기준일 : 2026년 8월 12일 연구 브리핑 제목 : Animal Behaviour Analysis Weekly Research Briefing 서비스 : Paper Quest (학술 검증 및 의사결정 지원 시스템)

1. 종합 선택 요약

구분논문 제목출판 / 버전교차검증 상태주요 특징 및 연구 위치
AI 최우선 추천Promptable Animal Pose Tracking Across SpeciesECCV 2026 CV4Ecology Workshop (accepted)VERIFIEDThe recommended paper aligns strongly with recent trends combining prompt engineering and animal pose estimation in ecology-related computer vision research, reflecting an emerging area that leverages cross-species generalization and promptability for improved tracking and behavior analysis.
사용자 최종 선택Promptable Animal Pose Tracking Across SpeciesECCV 2026 CV4Ecology Workshop (accepted)VERIFIED이번 주 읽기 지정 논문

2. 브리핑 추출 및 검증 분류 현황

  • 추출된 논문 후보 : 총 2편 (검증 완료: 2편, 평가 제외: 0편)
  • 데이터셋 : 2개 (APTv2, TigDog)
  • GitHub 도구/레포 : 2개 (SLEAP 1.5+, APTv2)
  • 평가 제외 항목 : Official GitHub code and pretrained weights for Promptable Animal Pose Tracking are currently unavailable and pending confirmation., Long-term drift stability on continuous multi-hundred-frame videos remains untested and requires future validation., Effectiveness of using promptable pose labels directly as pseudo labels for re-training SLEAP/DeepLabCut has not been demonstrated yet.
  • 불확실성 현황 :
    • 사실 검증 필요 (Fact Verification): 4건
    • 평가 근거 부족 (Insufficient Evidence): 4건
    • 연구 Open Question: 5건

3. 후보 논문 종합 평가 비교표

논문명교차검증성능 경쟁력방법론 신규성연구 흐름학술 유의미성실무·연구 적용재현 가능성출판 신뢰도최신성
Promptable Animal Pose Tracking Across Species 🌟[AI&최종선택]VERIFIED추가확인4점4점3점추가확인추가확인5점5점
Comparison of Multiple Video Tracking-Based Behavioral Summary Approaches for Compound DiscriminationVERIFIED추가확인3점3점3점추가확인추가확인3점5점

4. AI 추천 근거 및 위치 분석

🤖 AI 추천 논문: Promptable Animal Pose Tracking Across Species

  • 추천 이유 : The paper arXiv:2608.04995 is both the overall academic leader and the weekly topic leader specifically addressing promptable animal pose tracking across species, which aligns directly with the core topic of this week’s briefing. Despite its reproducibility status marked as NOT_VERIFIED and absence of available code and data, it holds academic acceptance in a relevant workshop (ECCV 2026 CV4Ecology Workshop), indicating peer recognition in the ecological and computer vision communities. The alternative candidate, bioRxiv 2026.07.20.739643v2, while related to behavioral summary approaches, is a preprint with reproducibility notes limited to PAPER_ONLY and does not directly cover the promptable tracking aspect across species. Therefore, arXiv:2608.04995 is recommended to prioritize relevance and academic standing for the user’s research focus.
  • 최근 연구 흐름상 위치 : The recommended paper aligns strongly with recent trends combining prompt engineering and animal pose estimation in ecology-related computer vision research, reflecting an emerging area that leverages cross-species generalization and promptability for improved tracking and behavior analysis.

주요 강점

  • Direct alignment with promptable animal pose tracking across species
  • Peer-reviewed acceptance at ECCV 2026 CV4Ecology Workshop
  • Focus on cross-species generalization relevant to ecological and behavioral research

주요 한계 및 위험요소

  • Lack of publicly available code and data limits reproducibility
  • Reproducibility not verified, potentially hindering immediate adoption
  • Uncertainty about empirical performance due to unavailable supporting resources

검증 필요 사항

  • ⚠️ Confirm reported results through independent replication once resources are available
  • ⚠️ Monitor for code and data release to enable reproducibility assessment
  • ⚠️ Evaluate generalizability to species or behaviors not covered in the original study

5. 선택 논문 상세 읽기 가이드 (Reading Checklist)

📖 최종 선택 논문: Promptable Animal Pose Tracking Across Species

  • 저자 : Le Li, Daniela Ivanova, Nicolas Pugeault (2026)
  • 출판/Preprint : ECCV 2026 CV4Ecology Workshop (accepted) (accepted)
  • 식별자/주소 : arXiv: 2608.04995
  • 코드 상태 : NOT_FOUND_AFTER_RETRIES (URL 미제공)
  • 데이터 상태 : NOT_FOUND_AFTER_RETRIES (URL 미제공)
  • 재현 가능성 : NOT_VERIFIED

🎯 읽을 때 집중해서 검토할 핵심 질문 (Key Questions)

  1. How does the method implement promptability across different animal species?
  2. What datasets and evaluation metrics were used for validation?
  3. Are there plans to release code or data to facilitate reproducibility?
  4. What are the limitations or failure cases noted by the authors?

🔬 후속 연구 및 아이디어 확장 질문 (Follow-up Research)

  • How can reproducibility be improved by providing code and data?
  • Can the approach be extended to more animal species or behaviors?
  • What comparative performance does the model achieve against other pose tracking methods?
  • Can integration with behavioral summary frameworks enhance ecological insights?

6. 각 후보 논문별 평가 근거 (Grounded Evidence) & 불확실성 분해

1. Promptable Animal Pose Tracking Across Species

  • 출판 정보 : ECCV 2026 CV4Ecology Workshop (accepted) (교차검증 상태: VERIFIED )
  • 6개 평가 축 점수 요약 :
    • 성능 경쟁력: 추가확인필요점 [상태: NEEDS_VERIFICATION] (No direct quantitative performance data, benchmark comparisons, or baseline results were found in the provided information or summary. Performance claims cannot be verified without access to detailed experimental results or official tables in the paper.)
    • 방법론적 신규성: 4점 [상태: SCORED] (The paper proposes a novel approach by using foundation-model dense features for user-prompted keypoint tracking in animal pose estimation, leveraging both supervised and unsupervised correspondence methods to achieve cross-species generalization. This method appears innovative in methodology compared to prior art which mostly dealt with species-specific pose estimation or tracking without foundation models.)
    • 연구 흐름 중요도: 4점 [상태: SCORED] (Foundation models and promptable approaches represent a cutting-edge trend in computer vision, particularly for tasks like pose tracking that require generalization beyond predefined classes or species. The paper aligns well with current trends towards foundation-model based generalist vision systems.)
    • 학술적 유의미성: 3점 [상태: SCORED] (The paper contributes a peer-reviewed workshop publication presenting a new methodological direction and extending foundation-model usage to ecological animal tracking. While impactful for academic interest, the significance is limited by lack of verified comparative evaluation or extensive community adoption yet.)
    • 실무/연구 적용 가치: 추가확인필요점 [상태: NEEDS_VERIFICATION] (Code and datasets are not found or publicly available currently, limiting immediate practical applicability or adoption by others in real experiments or applications.)
    • 재현 가능성: 추가확인필요점 [상태: NEEDS_VERIFICATION] (No code, datasets, or checkpoints have been found or verified available, and no reported reproduction attempts exist. Reproducibility is currently unverified and presumably low.)

📊 비교 연구 모듈

  • SOTA 및 상대적 위치 : No identified direct comparisons available with exact matching task, dataset, and metrics. The existing candidate is related but addresses a different task and evaluation context, placing the target paper as a novel contribution in foundation-model-based animal pose tracking across species.
  • 1) 직접 비교 연구 (Direct Comparison) :
    • 해당 동일 조건 직접 비교 연구 없음
  • 2) 유사 과업 연구 (Near Task Comparison / 정량 직접 비교 유예) :
    • Comparison of Multiple Video Tracking-Based Behavioral Summary Approaches for Compound Discrimination (2026): 과업 [Video tracking-based behavioral summary for compound discrimination] | 데이터셋 [상이한 벤치마크] | 지표 [상이한 평가 지표] | 유예 사유: The task focuses on behavioral summary using video tracking in compound discrimination, which is related but not the same as pose tracking across species. Datasets and metrics differ or are unspecified, disallowing direct performance comparison.
  • 3) 맥락상 관련 연구 (Contextual Related) :
    • 해당 맥락 관련 연구 없음
  • 4) 대표 선행 연구 (Representative Prior) :
    • 해당 대표 선행 연구 없음

📌 불확실성 3분할 (Uncertainty Breakdown)

  • 1) 사실 검증 필요 (Fact Verification) :
    • 🔍 Performance claims lack direct quantitative benchmark comparison evidence.
    • 🔍 Code and dataset for reproduction are not found and unverified.
  • 2) 평가 근거 부족 (Insufficient Evidence) :
    • ⚠️ No detailed experimental results or tables extracted to verify accuracy or SOTA claims.
    • ⚠️ No external benchmark or third-party reproduction evidence found.
  • 3) 연구적 Open Question :
    • ❓ How does the method quantitatively compare to existing animal pose estimation or tracking approaches on standardized datasets?
    • ❓ What is the robustness and generality performance across diverse species beyond reported cases?
    • ❓ Will released code and data be provided to enable reproduction and practical application?

📌 근거 삼분할 (Grounded Evidence)

  • 논문 원문 근거 :
    • 원문 스니펫에서 직접 추출되지 않음 (원문 수동 검증 필요)
  • 외부 출처 근거 :
    • 외부 기록 확인
  • AI 종합 해석 :
    • AI 종합 분석

2. Comparison of Multiple Video Tracking-Based Behavioral Summary Approaches for Compound Discrimination

  • 출판 정보 : bioRxiv preprint (교차검증 상태: VERIFIED )
  • 6개 평가 축 점수 요약 :
    • 성능 경쟁력: 추가확인필요점 [상태: NEEDS_VERIFICATION] (Paper compares multiple video tracking-based behavioral summary approaches but lacks direct quantitative benchmark comparisons against established state-of-the-art methods or official baselines; no verified performance table or external benchmark evidence available.)
    • 방법론적 신규성: 3점 [상태: SCORED] (The paper proposes a comparative study of multiple behavioral summary approaches for compound discrimination, including complex unsupervised representations and simpler tracking summaries, which represents a methodological contribution by analyzing these approaches in a novel comparative context.)
    • 연구 흐름 중요도: 3점 [상태: SCORED] (Video tracking and behavioral analysis for compound discrimination is relevant to ongoing trends in computational ethology and behavioral phenotyping, albeit not at the forefront of hot topics like deep learning pose estimation; it contributes to a meaningful emerging research area.)
    • 학술적 유의미성: 3점 [상태: SCORED] (Provides a comparative analysis in behavioral tracking for compound discrimination, which adds to academic knowledge and methodology in behavioral neuroscience and bioinformatics despite being a preprint and not peer-reviewed.)
    • 실무/연구 적용 가치: 추가확인필요점 [상태: NEEDS_VERIFICATION] (No code or data found after retries; practical applicability and usability depend on resource availability, which is currently missing or unverified.)
    • 재현 가능성: 추가확인필요점 [상태: NEEDS_VERIFICATION] (Code and data are not available; no evidence of third-party reproduction or detailed documentation to support reproducibility; reproducibility level is PAPER_ONLY.)

📊 비교 연구 모듈

  • SOTA 및 상대적 위치 : No direct state-of-the-art comparison studies found. Existing candidate paper addresses related but distinct tasks (animal pose tracking vs. behavioral summary for compound discrimination).
  • 1) 직접 비교 연구 (Direct Comparison) :
    • 해당 동일 조건 직접 비교 연구 없음
  • 2) 유사 과업 연구 (Near Task Comparison / 정량 직접 비교 유예) :
    • Promptable Animal Pose Tracking Across Species (2026): 과업 [Animal pose tracking across species] | 데이터셋 [상이한 벤치마크] | 지표 [상이한 평가 지표] | 유예 사유: Focuses on animal pose tracking rather than behavioral summary approaches for treatment discrimination; different task and metrics so direct comparison is not possible.
  • 3) 맥락상 관련 연구 (Contextual Related) :
    • 해당 맥락 관련 연구 없음
  • 4) 대표 선행 연구 (Representative Prior) :
    • 해당 대표 선행 연구 없음

📌 불확실성 3분할 (Uncertainty Breakdown)

  • 1) 사실 검증 필요 (Fact Verification) :
    • 🔍 Performance claims not directly verified with benchmark tables or official baselines.
    • 🔍 Code and data for reproducibility not found after retries.
  • 2) 평가 근거 부족 (Insufficient Evidence) :
    • ⚠️ Quantitative performance details and benchmarks are insufficient or absent.
    • ⚠️ No third-party reproduction or comprehensive documentation available.
  • 3) 연구적 Open Question :
    • ❓ How do the compared behavioral summary approaches rank quantitatively on standardized benchmarks?
    • ❓ Could releasing code and data improve practical adoption and reproducibility?

📌 근거 삼분할 (Grounded Evidence)

  • 논문 원문 근거 :
    • [Document Analysis] Comparison of multiple video tracking-based behavioral summary approaches, but quantitative results and direct benchmark comparisons are not detailed in snippet or summary.
  • 외부 출처 근거 :
    • 외부 기록 확인
  • AI 종합 해석 :
    • AI 종합 분석
Generated by Paper Quest - Evidence-based AI Academic Research Support System