Hacker Times
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
zargon
3 months ago
|
parent
|
context
|
favorite
| on:
Accelerating Gemma 4: faster inference with multi-...
They're using the term speculative decoding but doing MTP. It's the same thing as Nemotron, but Google removed the MTP heads from the original safetensora release. (They were not removed from the LiteRM format.)
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: