Differentiable Algebra Discovery Accelerates Grokking via a Non-Fourier Mechanism

저자: Bryan Cheng | 날짜: 2026 | URL: https://openreview.net/forum?id=CInhNzUwfZ 📄 PDF


⚠️ 이 페이지의 요약·평가·해설은 생성형 AI(Claude)가 자동 생성한 2차적 분석물입니다. 논문 원문의 저작권은 원저작자에게 있으며, 정확한 내용은 원문(위 DOI·arXiv 등 출처)을 확인하세요.

라이선스: OpenReview 공개(오픈액세스)

Essence

Figure 1

Figure 1. Left: Grokking curves for Z/97Z modular addition (mean over 3 seeds). FORGE + Grokfast (λ=7, blue) grokks at 1

FORGE는 rank-R bilinear product를 associativity, identity, inverse 3가지 algebra loss로 공동 학습시켜 group multiplication tensor의 CP factorization을 명시적으로 발견하게 하는 architecture로, 임의의 유한군에 대해 grokking을 비Fourier 메커니즘을 통해 극적으로 가속시킨다.

Motivation

Achievement

Figure 4

Figure 4. Grokking curves on four non-abelian groups (mean over 3 seeds, clipped at 15,000 steps). FORGE (blue) grokks w

  1. grokking 가속 성능: Z/97Z에서 FORGE+Grokfast가 10.20× 속도향상(MLP+Grokfast는 오히려 0.86×로 느려짐)을 달성했고, S3·D4·A4에서는 매칭된 MLP가 9개 런 전부 val=0.000으로 실패한 반면 FORGE는 ~10³ step 내 grokking에 성공했다.
  2. 비Fourier 메커니즘 발견: 6개 seed에 걸쳐 FORGE는 MLP와 동일한 정확도를 달성하면서도 sparse Fourier route가 아닌 질적으로 다른 경로(effective harmonic modes ≈23 vs ≈16, sign test p<0.016)를 사용함을 규명했다.
  3. 보편적 non-abelian grokking: abelian, dihedral, alternating, symmetric, quaternionic 등 A7(order 2,520), Z1009(order 1,009)를 포함한 모든 테스트 유한군에서 grokking에 성공했으며, 최소 비가해군(non-solvable group)인 A5에서 MLP 대비 10.1×, MLP+Grokfast 대비 3.0× 속도향상을 보였다.
  4. 보편 스케일링 법칙: grokking 시간이 732·|G|^0.170 (R²=0.785, 20개 군, 68 seed)이라는 sub-linear power law를 따름을 실증하여, 군의 order가 168배 증가해도 step 수는 2.4배만 증가함을 보였다.
  5. 이론적 특성화: Strassen-refined rank bound(rank_CP(T_{S4})≤55<64), mechanistic Wedderburn recovery(1,639개 conjugacy-class multiplicity 중 약 1,630개 정확히 복원), causal axiom-emergence theorem(identity axiom이 일반화보다 583±80 step 앞서 발생, tautology hypothesis 반박) 등 6개 명제를 제시했다.
  6. 도메인 일반화: Hamiltonian 시스템에서 invariant drift를 20–150,000× 감소시키고, QM9 U0 MAE를 15.3% 개선하여 group theory를 넘어선 general-purpose inductive bias임을 입증했다.

How

Figure 4

Figure 4. Grokking curves on four non-abelian groups (mean over 3 seeds, clipped at 15,000 steps). FORGE (blue) grokks w

Originality

Limitation & Further Study

Evaluation

Novelty: 5/5 Technical Soundness: 4/5 Significance: 4/5 Clarity: 4/5 Overall: 4/5

총평: grokking mechanistic interpretability를 순환군을 넘어 임의의 유한군으로 확장하고 이를 CP factorization이라는 엄밀한 수학적 틀로 정식화한 뛰어난 연구로, 이론과 실증을 균형있게 결합했다는 점에서 높은 평가를 받을 만하지만, 재현성과 일반화 범위에 대한 추가 검증이 필요하다.

같이 보면 좋은 논문

기반 연구SPECTER2 유사도 0.90로 LLM Reasoning and Safety Benchmarks와 Molecular Simulation and Generative Modeling가 맞닿아, 'Efficient and Equivariant Graph Networks for Predicting Quantum Hamiltonian'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구group theory 기반 신경망 학습의 이론적 기초를 제공한다.
다른 접근Transformer 기반 policy/value 모델을 이용한 조합 최적화 탐색이라는 유사한 방법론을 공유함.
기반 연구SPECTER2 유사도 0.90로 LLM Reasoning and Safety Benchmarks와 Molecular Simulation and Generative Modeling가 맞닿아, 'Equivariant Evidential Deep Learning for Interatomic Potentials'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
기반 연구SPECTER2 유사도 0.89 기준으로 'Differentiable Algebra Discovery Accelerates Grokking via a Non-Fourier Mechanism'의 AI4S 방법론을 'Extending the range of graph neural networks with global encodings'의 과학 생산·평가 맥락과 함께 보면 연구 자동화의 의미를 입체적으로 볼 수 있다.
기반 연구SPECTER2 유사도 0.89로 LLM Reasoning and Safety Benchmarks와 Molecular Simulation and Generative Modeling가 맞닿아, 'A Systematic Survey and Benchmark of Deep Learning for Molecular Property Prediction in the Foundation Model Era'가 이 ICML 2026 논문의 배경·대안·응용 맥락을 보완한다.
다른 접근grokking 현상을 다루는 다른 메커니즘 해석적 접근법이다.
후속 연구algebra discovery 관련 방법론을 확장하는 연구이다.
← 목록으로 돌아가기

🎧 Audio Overview

이 논문 리뷰를 팟캐스트형 오디오로 생성합니다. (Gemini · 키는 브라우저에만 저장 · 완성본은 이메일로도 전송)
▸ 고급: 구성 방향(대본 작성 지침) 직접 수정
속도 1.0x
⬇ MP3 다운로드