~ / skills / software-engineering / optimizing-attention-flash

optimizing-attention-flash

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens)…

Industry
software-engineering
License
Unverified
Source repo
NousResearch/hermes-agent · ★ 214,858
Source file
optional-skills/mlops/flash-attention/SKILL.md
View full SKILL.md on GitHub →

不会安装?看中文图文教程 →