← 返回资源分享
Mistral 7B
论文
论文
发布时间2023-10-10
发表arXiv:2310.06825
作者:Thibaut Lavril,Devendra Singh Chaplot,Guillaume Lample,Alexandre Sablayrolles,Marie-Anne Lachaux,Lucile Saulnier,Teven Le Scao,Thomas Wang,Arthur Mensch,Diego de Las Casas,Albert Q. Jiang,Timothée Lacroix,Chris Bamford,Florian Bressand,Gianna Lengyel,Lélio Renard Lavaud,Pierre Stock,William El Sayed
详细介绍
We introduce Mistral 7B v0.1, a 7-billion-parameter language model engineered for superior performance and efficiency. Mistral 7B outperforms Llama 2 13B across all evaluated benchmarks, and Llama 1 34B in reasoning, mathematics, and code generation. Our model leverages grouped-query attention (GQA) for faster inference, coupled with sliding window attention (SWA) to effectively handle sequences of arbitrary length with a reduced inference cost. We also provide a model fine-tuned to follow instructions, Mistral 7B -- Instruct, that surpasses the Llama 2 13B -- Chat model both on human and automated benchmarks. Our models are released under the Apache 2.0 license.
代码仓库 (7)
skypilot-org/skypilot官方PyTorch
mistralai/mistral-srcPyTorch
pwc-1/Paper-9/tree/main/2/mistralMindSpore
facebookresearch/fairseq2PyTorch
knowlab/bi-weekly-paper-presentation
ninglab/ecellmPyTorch
mgmalek/efficient_cross_entropyPyTorch
