常见开源大模型架构参数对照表(用于权重 / KV-Cache显存估算)

UpHub AI本地部署软件 ERP/进销存/订单管理软件 客户关系管理软件 网站GEO/SEO分析工具 上海哲涛科技 | 返回算力估算器

字段说明:

模型名称 X参数量(B) L层数 hidden_size num_ah
(Q头数)
num_kvh
(KV头数)
head_dim d_kv
KV头总维度
ctx_max
最大上下文
架构备注 官方网址
DeepSeek-V4-Flash 284 43 4096 64 1 512 512 1048576 DeepseekV4ForCausalLM https://github.com/deepseek-ai/
DeepSeek-V4-Pro 1600 61 7168 128 1 512 512 1048576 DeepseekV4ForCausalLM https://github.com/deepseek-ai/
DeepSeek-V4-Flash-0731 284 43 4096 64 1 512 512 1048576 DeepseekV4ForCausalLM https://github.com/deepseek-ai/
DeepSeek-V4-Flash-DSpark 284 43 4096 64 1 512 512 1048576 DeepseekV4ForCausalLM https://github.com/deepseek-ai/
DeepSeek-V4-Flash-Base 284 43 4096 64 1 512 512 1048576 DeepseekV4ForCausalLM https://github.com/deepseek-ai/
DeepSeek-V4-Pro-DSpark 1600 61 7168 128 1 512 512 1048576 DeepseekV4ForCausalLM https://github.com/deepseek-ai/
DeepSeek-V4-Pro-Base 1600 61 7168 128 1 512 512 1048576 DeepseekV4ForCausalLM https://github.com/deepseek-ai/
DeepSeek-V4-Pro-0813 1600 61 7168 128 1 512 512 1048576 DeepseekV4ForCausalLM https://github.com/deepseek-ai/
gemma-4-26B-A4B-it 26 30 2816 16 8 256 2048 262144 https://ai.google.com/
gemma-4-31B-it 31 60 5376 32 16 256 4096 262144 https://ai.google.com/
gemma-4-E4B-it 4 42 2560 8 2 256 512 131072 https://ai.google.com/
gemma-4-E2B-it 2 35 1536 8 1 256 256 131072 https://ai.google.com/
gemma-4-12B-it 12 48 3840 16 8 256 2048 262144 https://ai.google.com/
gemma-4-31B 31 60 5376 32 16 256 4096 262144 https://ai.google.com/
gemma-4-E4B 4 42 2560 8 2 256 512 131072 https://ai.google.com/
gemma-4-26B-A4B 26 30 2816 16 8 256 2048 262144 https://ai.google.com/
gemma-4-12B-it-assistant 12 4 1024 16 8 256 2048 262144 https://ai.google.com/
gemma-3-27b-it 27 62 5376 32 16 128 2048 262144 https://ai.google.com/
gemma-3-12b-it 12 48 3840 16 8 240 1920 262144 https://ai.google.com/
gemma-3-270m 0.27 18 640 4 1 256 256 32768 Gemma3ForCausalLM https://ai.google.com/
gemma-3-1b-it 1 26 1152 4 1 256 256 32768 Gemma3ForCausalLM https://ai.google.com/
gemma-3-1b-pt 1 26 1152 4 1 256 256 32768 Gemma3ForCausalLM https://ai.google.com/
Qwen3.5‑2B 2 24 2048 8 2 256 512 262144 GQA https://github.com/QwenLM/Qwen3.5
Qwen3.5‑4B 4 32 2560 16 4 256 512 262144 GQA https://github.com/QwenLM/Qwen3.5
Qwen3.5‑9B 9 32 4096 16 4 256 512 262144 GQA https://github.com/QwenLM/Qwen3.5
Qwen3.5‑27B 27 64 5120 24 4 256 1024 262144 GQA https://github.com/QwenLM/Qwen3.5
Qwen3.5‑35B-A3B 35 40 2048 16 2 256 512 262144 MOE https://github.com/QwenLM/Qwen3.5
Qwen3.5‑122B-A10B 122 48 3072 32 2 256 512 262144 MOE https://github.com/QwenLM/Qwen3.5
Qwen3.6‑27B 27 64 5120 24 4 256 1024 262144 GQA https://github.com/QwenLM/Qwen3.6
Qwen3.6‑35B-A3B 35 40 2048 16 2 256 512 262144 MOE https://github.com/QwenLM/Qwen3.6
Qwen3.8‑27B 27 64 5120 24 4 256 1024 262144 GQA https://github.com/QwenLM/Qwen3.8
Qwen3.8-2.4T-A95B 2400 92 8192 64 4 256 1024 262144 MOE https://github.com/QwenLM/Qwen3.8

显存估算提示