AI-Native DevOps: Building Self-Healing Infrastructure with Prometheus and LLMs
The pager rings at 3 AM. A Kubernetes pod is crash-looping. Traditionally, you wake up, log in, run `kubectl logs`, and spend an hour debugging. In the era of AI-Native DevOps, this workflow is obsolete. An AI agent has already intercepted the Prometheus alert, analyzed the stack trace, cross-referenced it with recent git commits, and even applied a rollback—all before you woke up. This guide shows you how to build this system.
Architecture: The Loop
The core of AI-Native DevOps is the OODA Loop (Observe, Orient, Decide, Act) automated by an LLM Agent.
The Components
- Sensor: Prometheus Alertmanager sends a webhook when a metric spikes.
- Brain: An LLM Agent (LangChain/AutoGPT) receives the webhook.
- Hands: The Agent has tools: `kubectl`, `git`, `aws-cli`.
- Judge: A Human-in-the-Loop (via Slack) approves high-risk actions.
Security: The Rogue Agent Problem
Giving an AI `kubectl admin` access is a recipe for disaster. If the LLM hallucinates, it could delete your production database.
Mitigation Strategies
Least Privilege (RBAC)
The Agent's ServiceAccount should only have permission to `get logs` and `restart deployment`. It should never have `delete persistentvolume`.
Human Approval
For any write operation (Act), the agent must send a message to Slack: "I want to restart service X. Approve?"
Building the Agent (Node.js)
This is a simplified webhook handler that acts as our AI DevOps Engineer.
import express from 'express';
import { ChatOpenAI } from "@langchain/openai";
import { initializeAgentExecutorWithOptions } from "langchain/agents";
import { Tool } from "langchain/tools";
import { exec } from 'child_process';
const app = express();
app.use(express.json());
// Define a Custom Tool for Kubectl
class KubectlLogsTool extends Tool {
name = "kubectl_logs";
description = "Get logs for a pod. Input should be pod name.";
async _call(podName: string) {
return new Promise((resolve) => {
exec(`kubectl logs ${podName} --tail=50`, (error, stdout) => {
resolve(stdout || "No logs found or error.");
});
});
}
}
app.post('/webhook/alertmanager', async (req, res) => {
const alert = req.body.alerts[0];
const podName = alert.labels.pod;
const model = new ChatOpenAI({ temperature: 0 });
const tools = [new KubectlLogsTool()];
const executor = await initializeAgentExecutorWithOptions(tools, model, {
agentType: "openai-functions",
});
const prompt = `
Alert received for pod ${podName}.
The error is: ${alert.annotations.description}.
1. Get the logs.
2. Analyze the root cause.
3. Suggest a fix.
`;
const result = await executor.run(prompt);
// Send result to Slack (omitted)
console.log("AI Investigation Result:", result);
res.sendStatus(200);
});Observability Hooks (Spring Boot)
For the AI to understand the application, the app must expose deep diagnostics. Standard Actuator health checks aren't enough. We need a "Diagnostic Endpoint" designed for LLMs.
@RestController
public class DiagnosticController {
@Autowired
private DataSource dataSource;
@GetMapping("/internal/diagnose")
public Map<String, Object> getDiagnostics() {
Map<String, Object> status = new HashMap<>();
// 1. DB Connection Pool Status
if (dataSource instanceof HikariDataSource) {
HikariPoolMXBean pool = ((HikariDataSource) dataSource).getHikariPoolMXBean();
status.put("db_active_connections", pool.getActiveConnections());
status.put("db_threads_awaiting", pool.getThreadsAwaitingConnection());
}
// 2. Memory Usage
status.put("heap_used", Runtime.getRuntime().totalMemory() - Runtime.getRuntime().freeMemory());
// 3. Last 5 Exceptions (from a memory buffer)
status.put("recent_errors", ErrorBuffer.getLatest());
return status;
}
}Interactive Incident Workflow
Alert Fired
Prometheus detects high latency.
Agent Investigates
Agent queries `/internal/diagnose` and sees "db_threads_awaiting: 50".
Hypothesis
Agent hypothesizes: "Connection pool exhaustion due to slow query."
Remediation
Agent scales up the ReplicaSet to handle load OR restarts the pod (if configured).
Conclusion
AI-Native DevOps is not about replacing SREs; it's about giving them superpowers. By offloading the "investigation" phase to an AI, humans can focus on the "architecture" phase, building systems that are resilient by design.