
概览
主要功能
- 硬件检测和分析
- 模型大小和量化推荐
- VRAM和RAM需求估算
- 每个变体的性能预期
- 支持多个Gemma 4版本
- CPU和GPU推理指导
价格
- 模型
- Free
- 分类
- 微为架的给布系统
- 评分
- 4.3 / 5 (6)
使用场景
为您的GPU选择合适的Gemma 4变体
开发者可以快速确定哪个Gemma 4大小和量化级别适合其可用的VRAM,避免本地推理期间的内存不足崩溃。
规划仅限CPU的推理设置
没有专用GPU的爱好者可以使用匹配器来查找一个Gemma 4变体,该变体可以在系统RAM和CPU上以可接受的性能运行,并具有现实的性能预期。
评估本地LLM的硬件升级
研究人员可以比较哪些Gemma 4版本在不同的VRAM或RAM层级上变得可访问,从而有助于证明本地模型工作的硬件投资是合理的。
平衡模型质量和速度
用户可以查看推荐的量化级别,以权衡输出质量与推理速度,从而选择最适合其工作流程的变体。
优点 & 缺点
优点
- 节省评估模型兼容性的时间
- 考虑有限硬件的量化选项
- 对初学者和高级用户都很有用
- 帮助避免内存不足错误
缺点
- 仅限于Gemma 4模型家族
- 推荐取决于准确的硬件检测
- 可能无法考虑每个运行时或后端
对决战绩
在万神殿中参与了 2 对决。
Last 2 battles
评测
6 个评分的平均值。
登录以留下评测。
Use it every day
Honestly didn't expect to like it this much. Support for multiple Gemma 4 versions is exactly what I needed, and useful for both beginners and advanced users. I do wish recommendations depend on accurate hardware detection, but I reach for it almost every day now and it just clicks.
Does the job
Pretty happy overall. Support for multiple Gemma 4 versions just works and useful for both beginners and advanced users. Recommendations depend on accurate hardware detection can be annoying, but no dealbreakers — I'd recommend it to a friend without hesitating.
Skeptical, then convinced
I went in skeptical — most tools in this space overpromise. It actually delivers on support for multiple Gemma 4 versions, and helps avoid out-of-memory failures caught me off guard. still, I'd recommend giving it a real trial.
Solid for our team
We rolled this out across the team last quarter and saves time evaluating model compatibility. Model size and quantization recommendations fits neatly into how we already work, and vRAM and RAM requirement estimates removed a step we used to do by hand. Recommendations depend on accurate hardware detection, which is the main caveat, but it has held up under daily use.
Solid for our team
We rolled this out across the team last quarter and useful for both beginners and advanced users. VRAM and RAM requirement estimates fits neatly into how we already work, and vRAM and RAM requirement estimates removed a step we used to do by hand. May not account for every runtime or backend, which is the main caveat, but it has held up under daily use.
Solid for our team
We rolled this out across the team last quarter and useful for both beginners and advanced users. Performance expectations per variant fits neatly into how we already work, and guidance for CPU and GPU inference removed a step we used to do by hand. Limited to the Gemma 4 model family, which is the main caveat, but it has held up under daily use.
问答
How to fix OOM errors when running Gemma 4 locally?
当使用本地环境运行Gemma 4出现OOM错误时,请尝试以下步骤:1) 降低上下文长度,使用--ctx-size 4096或更低的值。2) 使用较低的量化,切换至Q3或Q2。3) 降低模型大小,使用26B MoE而不是31B,或E4B而不是26B。4) 降低并行程度,使用OLLAMA_NUM_PARALLEL=1与OLLAMA搭配使用。Gemma 4的KV缓存因为256K的上下文窗口而特别大上, 所以上下文长度是占用Vram的最大因素。
Asked by Elif Yildiz · Feb 15, 2026
Can I run Gemma 4 on my phone?
是的。请在安卓或iOS上安装Google AI Edge Gallery,然后下载Gemma 4 E2B模型 (5B 参数,~3 GB)。该模型可以完全离线运行,不需要API密钥。E4B模型 (9B参数)可能在10+GB内存的手机上工作,但在内存较少的设备上可能会崩溃。推荐在手机上使用E2B
Asked by Celia Ramirez · Jan 29, 2026
如何在 LM Studio 中运行 Gemma 4?
打开 LM Studio,转到搜索标签,并键入“gemma 4”。您将在各种量化水平中找到 E2B、E4B、26B MoE 和 31B Dense 模型等级的 GGUF 文件。LM Studio会自动突出显示哪些版本符合您的可用 VRAM。单击下载,然后切换到聊天标签开始说话。LM Studio还支持 NVIDIA GPU 的 EXL2 格式。
Asked by Yuki Kobayashi · Jan 13, 2026
What's the difference between Gemma 4 26B MoE and 31B Dense?
Gemma 4 26B MoE (混合权重专家)总参数数量为 27B,但每个token仅激活 ~4B。因此,它的速度远快于实际大小——与4B模型的速度相当——同时保持较高质量。它可以放入 12–16 GB VRAM,是消费级硬件 (RTX 4070 Ti, MacBook Pro 16 GB) 的最佳选择。31B Dense 在每个token时激活所有 31B 参数,输出质量最高,但必须有20+ GB VRAM。它适用于工作站 (RTX 4090, M-series 32 GB+)
Asked by Marisol Pena · Jan 13, 2026
点击 GEMMA 4 或下一分 EXL2 入子?
秀。 GEMMA 4 或子、(excluding 寸B射 Dense) 组有分 EXL2入子为其中国为成細。 EXL2 文件上入券三入子 在 HuggingFace。存在一分 EXLlamaV2 给券算了、 EXL2 号使武一代 断起史,下给一分分帔成为使用 VRAM/名了不滤成为券。。。
Asked by Gabriel Duarte · Jan 6, 2026
提问
微为架的给布系统 的替代品

高性能LLM网关,统一1000+模型于单一API

下一代深度寻求的推理聚焦的 AI 模型

开源的 Mixture-of-Experts 模型,提供相当于 GPT-4o 水平的推理能力,且成本仅为其一小部分。

xAI开发的会话式 AI 工具,旨在提高推理、研究和实时问答能力。

Meta开源的多语言开放权重LLM,专为高效文本生成提供了强大的支持。

音频转文字工具 AI 算法快速生成文本稿

一款开源大型语言模型,擅长推理、数学和编码任务,采用 MIT 许可证,可免费使用和修改。

为当前入学式的广頃场橋能纳录。




