Why Local Agentic Coding Matters
Cloud LLMs are great, but every prompt ships your source code to someone else's GPU. For teams working on proprietary auth modules, billing logic, or regulated data, that's a non-starter. The Antigravity SDK now supports local model execution with initial backing for Gemma 4 26B A4B through Google AI Edge's LiteRT runtime — meaning your agent orchestrates file edits, runs tests, and audits code without a single byte leaving your machine.
The trade-offs are real: you need >24GB VRAM or unified memory, and local inference is slower than cloud. But for privacy-sensitive workflows, the trade is worth it. If you want to understand how context is engineered for agents running in the background, our context engineering deep dive for background coding agents is a good companion read.
In this tutorial, you'll go from zero to a working offline agent that can inspect files, write scripts, and run them — all on your laptop.

Step 1: Environment and Model Import
Start with a clean virtual environment. The litert-lm CLI handles the model download directly from Hugging Face.
# Create and activate an isolated environment
python3 -m venv .venv
source .venv/bin/activate
# Install the Antigravity SDK and LiteRT-LM runtime
pip install google-antigravity litert-lm
# Import Gemma 4 26B A4B (GPU-optimized build)
litert-lm import \
--from-huggingface-repo=litert-community/gemma-4-26B-A4B-it-litert-lm \
gemma-4-26B-A4B-it-gpu.litertlm \
gemma4-26b
Step 2: Your First Offline Agent
Create agy_sample.py and paste the following. Note the lightweight() config — it trims the agent's toolset for faster local startup.
import asyncio
import os
from google.antigravity import Agent, LiteRTAgentConfig
from google.antigravity.hooks import policy
# Point to the model you imported in Step 1
MODEL_PATH = os.path.expanduser("~/.litert-lm/models/gemma4-26b/model.litertlm")
async def main():
print(f"Using local LiteRT model: {MODEL_PATH}. Local inference may take several minutes.")
config = LiteRTAgentConfig(model_path=MODEL_PATH).lightweight()
async with Agent(config) as agent:
response = await agent.chat("What files are in the current directory?")
async for token in response:
print(token, end="", flush=True)
if __name__ == "__main__":
asyncio.run(main())
Run it with python agy_sample.py. First inference will be slow — the model has to load into VRAM — but subsequent calls are dramatically faster.
Step 3: Let the Agent Build Its Own Tool
The real power shows up when you give the agent a workspace and let it write files. This example asks Gemma 4 26B to build a live terminal resource monitor using psutil and rich, generate a requirements.txt, and test the result.
import asyncio
import os
from google.antigravity import Agent, LiteRTAgentConfig
from google.antigravity.hooks import policy
PROMPT = (
"Build a CLI tool using psutil and rich that displays a live-updating "
"terminal dashboard showing CPU usage, memory consumption, and the top 5 "
"memory-intensive processes. Save as 'monitor.py', create a requirements.txt, "
"and test that it works."
)
MODEL_PATH = os.path.expanduser("~/.litert-lm/models/gemma4-26b/model.litertlm")
WORKING_DIR = os.path.expanduser("~/agy-test")
os.makedirs(WORKING_DIR, exist_ok=True)
os.chdir(WORKING_DIR)
async def main():
print(f"Using local LiteRT model: {MODEL_PATH}. Please wait for local inference.")
config = LiteRTAgentConfig(
model_path=MODEL_PATH,
workspaces=[WORKING_DIR],
policies=[policy.allow_all()],
).lightweight()
async with Agent(config) as agent:
response = await agent.chat(PROMPT)
async for token in response:
print(token, end="", flush=True)
if __name__ == "__main__":
asyncio.run(main())
Step 4: The Architect-Builder Pattern (Hybrid Mode)
You don't have to choose between cloud and local. A pattern that works well in practice is using a cloud model as the architect (planner, conductor) while a local swarm of Gemma 4 26B instances does the heavy lifting. This keeps sensitive code on-device while still benefiting from cloud-scale reasoning for planning.
A concrete example: auditing and patching three vulnerable modules (auth.py, billing.py, database.py). The cloud architect plans the audit; the local swarm reads each file, proposes patches, and runs tests — all without the source ever leaving the machine. Token costs drop dramatically because the expensive cloud model only sees plans, not file contents.
Step 5: Plug-and-Play with OpenAI-Compatible Backends
If you already run Ollama, LM Studio, or vLLM, you can swap backends without touching your agent code. Use LocalOpenAIAgentConfig instead of LiteRTAgentConfig and point it at your local server. Your tools, hooks, and orchestration stay identical — only the inference engine changes.

Limitations and Caveats
- Hardware floor is real. 24GB VRAM is a lot. On a 16GB laptop, you'll be swapping and the experience will be painful. Consider smaller quantized models or the OpenAI-compatible path with a remote server.
- Cold start latency. Loading a 26B model into VRAM takes minutes. Keep the process warm for iterative work, or accept the startup tax.
- Tool reliability degrades. Smaller local models are more prone to malformed tool calls than frontier cloud models. Always wrap
allow_all()policies with a sandboxed workspace — never point the agent at your home directory. - Context windows are tighter. Long sessions will hit limits faster than with cloud models. Chunk your tasks.
- No free lunch on quality. Gemma 4 26B is impressive for its size, but it is not Gemini 3.8 Flash. Reserve local execution for privacy-critical or cost-sensitive jobs.
Where to Go Next
- Measure your throughput. Benchmark tokens/sec on your hardware before committing to a workflow. If it's under ~15 tok/s, hybrid Architect-Builder is the only sane pattern.
- Build a tool library. The agent is only as good as the tools you expose. Start with filesystem, shell, and test runners — then add domain-specific tools.
- Read the SDK README for the full list of hooks and policies, and file issues on the GitHub tracker when something breaks.
- Study systems-level optimization. If you're pushing inference throughput on accelerators, our systems engineering playbook for serving large MoE models at scale covers the hardware-side tricks that translate directly to local inference.

The Bottom Line
The Antigravity SDK's local model support isn't a toy. With Gemma 4 26B A4B via LiteRT, you get a genuinely useful offline agent that can audit code, write scripts, and run tests without leaking a single line to the cloud. The Architect-Builder pattern lets you keep the smart parts in the cloud and the sensitive parts on-device.
Start with the sample script, then graduate to the file-writing example. Once you're comfortable, wire it into your real workflow — and keep the workspace sandboxed.
Sources: Introducing support for local AI models in the Antigravity SDK
Further reading: