llama.cpp(开源本地大模型推理引擎)
官方介绍
llama.cpp(ggml-org)是最主流的开源大模型本地推理引擎:纯 C / C++ 实现,以「最少配置、广泛硬件上的先进性能」为目标,定义了 GGUF 量化文件格式。CPU、Apple Silicon(Metal)、NVIDIA / AMD GPU 全覆盖,是 Ollama、LM Studio 等本地工具的底层。
核心功能
- GGUF 量化格式 + 转换脚本(convert_*.py),4-bit / 2-bit 等量化让小内存设备也能跑大模型
- 跨平台后端:CPU / Metal / CUDA / Vulkan 等
llama-server:OpenAI 兼容本地 API 服务,内置 Web UI-hf参数直接从 Hugging Face 下载并运行模型
使用说明
- 安装:从源码构建(见 build.md),或用 Homebrew 等包管理器
- 运行:
llama cli -hf <user>/<model>[:quant]下载即跑;llama serve -hf <model>起 OpenAI 兼容服务 - 模型需为 GGUF 格式,其他格式用仓库内脚本转换
官方链接
- GitHub 仓库:github.com/ggml-org/llama.cpp
- 构建文档:docs/build.md
- GGUF 格式说明:ggml/docs/gguf.md