Search for a command to run...
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
What can I help you find?
Datasets, papers, notebooks and GPUs