现代 LLM 的 Memory-Compute Trade-off:从 MHA 到 MLA、Sparse 与 Linear Attention2026年9月11日·技术杂谈·LLMsAttentionMLADeepSeek长上下文