Streaming chat
Ask for explanations, refactors, or debug help and watch Llama 3.3 70B answer token by token over SSE.
AI code assistant · live
Codistant
Stream answers from Llama 3.3 70B, keep conversation memory, and run multi-step reviews - on Cloudflare, with the UI on Vercel.
Three capabilities, one assistant.
Ask for explanations, refactors, or debug help and watch Llama 3.3 70B answer token by token over SSE.
Durable Objects with SQLite keep conversation history and project context across reloads - per session.
Cloudflare Workflows run a durable three-step review: structure, issues, then concrete improvements.
A hybrid edge stack with a clear split of responsibilities.
Static UI on Vercel at codistant.v-ai.org - landing plus the chat app.
API router on Cloudflare Workers handles chat, history, context, and review.
Durable Objects store memory; Workflows run the multi-step code review.
Llama 3.3 70B generates the streaming answers and review steps.
Shipping Codistant forced real edge-AI tradeoffs - not just a chat demo.
SSE from a Worker to a Vercel origin needs CORS on every response - including the streamed body - or the browser quietly fails mid-flight.
Durable Objects with embedded SQLite made per-session history and project context feel local, without standing up a separate database.
Code review is three durable steps with retries. Orchestrating that as a Workflow beat stuffing everything into one fragile request.
UI on Vercel, intelligence on Cloudflare. Each platform does what it’s good at - and the config surface stays small.
Open to opportunities
I build AI-native products end to end: edge runtimes, streaming UX, and practical agent workflows. If you’re hiring for AI / full-stack / platform roles, let’s talk.