Instructions to use google/gemma-4-E2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use google/gemma-4-E2B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("google/gemma-4-E2B") model = AutoModelForMultimodalLM.from_pretrained("google/gemma-4-E2B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Gemma4 grew after update?
So, I used to be able to run gemma4:e2b comfortably in 8 GB memory on my Mac. However, today I installed an update of the model and... it doesn't fit any more! I use ollama, is it possible to revert to previous version of the model?
500 Internal Server Error: model requires 6.9 GiB but only 4.8 GiB are available (after 512.0 MiB overhead)
Hi @jjido
It may be possible if the previous model files are still on disk. Pulling gemma4:e2b again won’t restore the old version, since the tag now points to the latest one. If you see two large files of different sizes, the older one may be the previous weights. If the old files are gone, it’s probably easier to reduce memory use instead of reverting. Try a smaller context length (OLLAMA_CONTEXT_LENGTH=4096), since the memory estimate includes the KV cache. Hope this helps .
Thanks
Thanks. I tried going from 8KB context to 4KB but I am getting the same message.
One thing I noticed, it is awfully specific about how much memory is available: it is always 4.8 GiB, no matter how many other apps are running on the computer. I wonder if there is something wrong with the download.