https://www.together.ai/blog/hungry-hungry-hippos-towards-language-modeling-with-state-space-models (opens in new tab)
--- title: Hungry Hungry Hippos: Towards language modeling with state space models --- ⚡️ FlashAttention-4: up to 1.3× faster than cuDNN on NVIDIA Blackwell → Introducing Together AI's new look → 🔎 ATLAS: runtime-learning accelerators delivering up to 4x faster LLM inference → ⚡ Together GPU Clusters: self-service NVIDIA GPUs, now generally available → 📦 Batch Inference API: Process billions of tokens at 50% lower cost for most models → 🪛 Fine-Tuning Platform Upgrades: Larger Models, Lo...
Read the original article