Topic information becomes decodable within the first few layers and stays decodable everywhere (kNN ≈ 0.9 from layer 2 on). Yet through the first half of the network the classes are visually entangled — they collapse onto overlapping manifolds (silhouette ≈ 0) — and only re-separate into clean clusters in the second half.