Searching...
A design space for controllable quench-and-relax architectures
Can we swap softmax attention for energy-based attention?