[D] Information geometry, anyone?
reddit.com·15h·
Flag this post

The last few months I’ve been doing a deep-dive into information geometry and I’ve really, thoroughly enjoyed it. Understanding models in higher-dimensions is nearly impossible (for me at least) without breaking them down this way. I used a Fisher information matrix approximation to “watch” a model train and then compared it to other models by measuring “alignment” via top-k FIM eigenvalues from the final, trained manifolds.

What resulted was, essentially, that task manifolds develop shared features in parameter space. I started using composites of the FIM top-k eigenvalues from separate models as initialization points for training (with noise perturbations to give GD room to work), and it positively impacted the models themselves to train faster, with better accuracy, and fewer a…

Similar Posts

Loading similar posts...