Serverless AI: Running Local LLMs in the Browser with WebGPU and WebLLM

📅 December 19, 2025⏱️ 28 min read🏷️ Web AI

The cloud is expensive. Inference costs for models like GPT-4 are a major bottleneck for scaling AI apps. But what if you could offload 100% of the compute to the user's device? Thanks to WebGPU and highly optimized libraries like WebLLM and Transformers.js, you can now run Llama 3 8B, Phi-3, and Mistral directly in Chrome, running at 50+ tokens/sec on a modern MacBook or gaming PC.

Why WebGPU Changes Everything

For years, running AI in the browser meant using WebGL, which was designed for graphics, not general-purpose compute (GPGPU). WebGPU is the successor, offering low-level access to the GPU with compute shaders that map perfectly to the matrix multiplication operations required by Transformers.

Before (WebGL)

High overhead, data marshaling bottlenecks. Llama 2 ran at 2-3 tokens/sec. Usable only for tiny models.

Now (WebGPU)

Direct compute shader access. Llama 3 8B runs at 40-50 tokens/sec on an M3 MacBook Air.

Security: The Privacy-First Advantage & New Risks

The biggest selling point of browser-based AI is Privacy. No data ever leaves the user's device. This is massive for HIPAA-compliant apps, legal tech, or personal journaling apps where users trust no one with their data.

But what about the risks?

Running models on the client introduces new vectors that backend engineers often ignore:

Model Theft

If you finetune a proprietary model and ship the weights to the client, a user can simply inspect the Network tab and download your `.onnx` or `.wasm` files. Rule: Never ship proprietary IP models to the client. Use open weights (Llama, Mistral) or distill knowledge into a generic model.

XSS & Prompt Injection (Client-Side)

If your client-side model processes untrusted input (e.g., from a shared document), it is still vulnerable to Prompt Injection. If the model output is rendered into the DOM without sanitization, it can lead to XSS.

Next.js Implementation: Transformers.js

We will use `transformers.js` to run a text generation pipeline in a Web Worker (to avoid freezing the main UI thread).

worker.js (Web Worker)
import { pipeline, env } from '@xenova/transformers';

// Skip local model checks
env.allowLocalModels = false;
env.useBrowserCache = true;

class TextGenerationPipeline {
  static task = 'text-generation';
  static model = 'Xenova/TinyLlama-1.1B-Chat-v1.0';
  static instance = null;

  static async getInstance(progress_callback = null) {
    if (this.instance === null) {
      this.instance = await pipeline(this.task, this.model, { progress_callback });
    }
    return this.instance;
  }
}

self.addEventListener('message', async (event) => {
  const { text } = event.data;

  const generator = await TextGenerationPipeline.getInstance(data => {
    self.postMessage({ status: 'progress', data });
  });

  const output = await generator(text, {
    max_new_tokens: 200,
    temperature: 0.7,
    callback_function: (beams) => {
        // Stream tokens back
        const decodedText = generator.tokenizer.decode(beams[0].output_token_ids, { skip_special_tokens: true });
        self.postMessage({ status: 'update', output: decodedText });
    }
  });

  self.postMessage({ status: 'complete', output: output[0].generated_text });
});

And the React component to consume it:

ChatComponent.tsx
'use client';
import { useEffect, useRef, useState } from 'react';

export default function ChatComponent() {
  const worker = useRef<Worker>(null);
  const [output, setOutput] = useState('');
  const [ready, setReady] = useState(false);

  useEffect(() => {
    if (!worker.current) {
      worker.current = new Worker(new URL('./worker.js', import.meta.url), {
        type: 'module',
      });

      worker.current.onmessage = (e) => {
        if (e.data.status === 'update') {
            setOutput(e.data.output);
        }
      };
    }
  }, []);

  const generate = () => {
    worker.current?.postMessage({ text: "Explain quantum computing to a 5 year old." });
  };

  return (
    <div className="p-4 bg-gray-900 rounded-xl border border-gray-800">
        <button onClick={generate} className="bg-purple-600 px-4 py-2 rounded text-white mb-4">
            Generate on Device
        </button>
        <div className="prose prose-invert">
            {output}
        </div>
    </div>
  );
}

Backend Coordination (Spring Boot)

Even in a "serverless" AI architecture, the backend still plays a crucial role: syncing state, managing user profiles, or offering a fallback to a larger server-side model if the user's device is too weak (e.g., an old mobile phone).

Java (Spring Boot) Device Capability Check
@RestController
@RequestMapping("/api/config")
public class ModelConfigController {

    @PostMapping("/negotiate")
    public ModelConfig negotiateModel(@RequestBody DeviceCapabilities device) {
        // If device has < 4GB VRAM or is mobile, suggest server-side
        if (device.getVram() < 4096 || device.isMobile()) {
            return new ModelConfig(ExecutionMode.SERVER, "gpt-4o-mini");
        }

        // Else suggest local Llama 3
        return new ModelConfig(ExecutionMode.LOCAL, "Llama-3-8B-Quantized");
    }
}

class DeviceCapabilities {
    private int vram;
    private boolean mobile;
    // getters/setters
}

Pro-Tips for Performance

Use Quantization

Always use 4-bit (q4f16) or 8-bit quantization. It reduces memory usage by 70% with negligible quality loss.

Cache Model Weights

The Cache API is your friend. Users shouldn't download 4GB every time they visit.

Web Workers

Never run inference on the main thread. It will lock up the UI and kill the scroll performance.

Conclusion

Local AI is the future of privacy-preserving, zero-cost interaction. By leveraging WebGPU and smart backend coordination, we can build apps that are both powerful and respectful of user data. The tools are here—it's time to build.

🌌
Purple Dream
Active Theme