Skip to content
https://abc.microfintool.com/

Mobile & Accessories

  • Welcome to ABC Tool: Your Ultimate Portal for Smartphones and Accessories
  • About / Contect
    • PRIVACY POLICY
  • Blog
Google’s latest trick gets Gemma 4 running 3x faster right on your phone

Google’s latest trick gets Gemma 4 running 3x faster right on your phone

Posted on May 6, 2026 By safdargal12 No Comments on Google’s latest trick gets Gemma 4 running 3x faster right on your phone
Blog

[ad_1]

TL;DR

  • Google has introduced new assistant models, called “drafters,” that could significantly speed up Gemma 4.
  • Drafters work by predicting sections of prompts to the main model, which can focus on processing them in bigger batches.
  • This allows the model to use the memory and the compute more efficiently.

Google’s recently launched Gemma 4 edge AI models are especially designed to run locally on consumer-hosted hardware. While favorable from a privacy standpoint, local models can easily hog resources and slow down results, rendering them ineffective. So, Google is now offering a potential solution, which it claims can speed up Gemma 4 models by up to three times.

Google recently released Multi-Token Prediction (MTP) drafters for Gemma 4. These drafters are essentially smaller, assistive models that help the primary model by “predicting” part of the user’s request. These smaller models also work in parallel to the main model to manage the compute more effectively.

Don’t want to miss the best from Android Authority?

google preferred source badge light@2x
google preferred source badge dark@2x

How does MTP improve Gemma 4?

The process uses a technique called “Speculative Decoding,” in which the drafter models predict upcoming words in the prompt even before the main Gemma model has read through it. While the drafter moves on to the next sequence of words, the main model verifies the predicted set of words at the same time.

If the model accepts the drafted version, it moves on to verify the next set. If it disagrees, it replaces the incorrect word or chunk.

While the extra work may sound counterintuitive, it’s actually not. Let me give you an oversimplified explanation of why MTP works.

The speed of processing is not just determined by the processing hardware (typically GPU cores) but by the memory bandwidth (VRAM). That’s because the model has to be referenced with each new request. So, by combining multiple words into a single chunk, the model must be referenced only once rather than multiple times, thus, shifting the load from the memory to the processing unit.

In addition to making these changes, Google says it is also working to optimize Gemma 4 models of different weights for specific hardware, such as the Apple Silicon or the popular Nvidia A100.

The MTP drafters for Gemma 4, alongside the primary model, can use platforms such as HuggingFace or Kaggle, tools like Ollama, or through Google’s own AI Edge Gallery on Android or iOS.

Thank you for being part of our community. Read our Comment Policy before posting.

[ad_2]

Source link

Post Views: 38
Tags: Google News

Post navigation

❮ Previous Post: A new leak says Apple’s all-screen iPhone may ditch buttons too
Next Post: The Boring Internet | Terry Godier ❯

You may also like

DuckDuckGo makes its ‘no-AI’ search engine easier to access as its traffic booms
Blog
DuckDuckGo makes its ‘no-AI’ search engine easier to access as its traffic booms
June 1, 2026
Tecno Pop X 5G's key features, design, and launch date revealed
Blog
Tecno Pop X 5G's key features, design, and launch date revealed
April 18, 2026
Price updates for apps, In-App Purchases, and subscriptions – Latest News
Blog
Price updates for apps, In-App Purchases, and subscriptions – Latest News
April 26, 2026
Are You Eligible to Claim Part of Apple’s 0M AI iPhone Settlement? How to Find Out
Blog
Are You Eligible to Claim Part of Apple’s $250M AI iPhone Settlement? How to Find Out
June 14, 2026

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Recent Posts

  • Fix Outlook for Mac 16.110 Missing Email History
  • Anthropic’s Claude Tag: Smarter Slack Assistant
  • Google Home Facial Recognition Update Boosts Accuracy
  • Meta Pauses Employee Tracking After Data Leak
  • US Accelerates Post-Quantum Cryptography Deadline to 2030

Recent Comments

  1. nszgpwrtqv on WhatsApp is now testing its subscription service, here's what you get and how much it costs
  2. qmkffkjqvx on Man dies covered in necrotic lesions after amoebas eat him alive
  3. Declan Chidlow on History of Game Console Web Browsers: Evolution & Tech
  4. ALEXAnync on Meta steals a tactic from Tesla and builds data centers in tents
  5. ALEXAnync on The Delivery You Didn’t Order: Breaking Down the ‘Free Phone’ Scam

Archives

  • June 2026
  • May 2026
  • April 2026

Categories

  • Blog

Copyright © 2026 Mobile & Accessories.

Theme: Oceanly News by ScriptsTown