AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-25 · America/Los_Angeles · 论文 · #8

Parts-of-Speech as Emergent Categories in SAE Latent Space

Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ…

Parts-of-Speech as Emergent Categories in SAE Latent Space

热度 38.6 / 100;排名与评分保留该期记录。