AI-Native DevOps: Building Self-Healing Infrastructure with Prometheus and LLMs

📅 December 21, 2025⏱️ 35 min read🏷️ DevOps AI

The pager rings at 3 AM. A Kubernetes pod is crash-looping. Traditionally, you wake up, log in, run `kubectl logs`, and spend an hour debugging. In the era of AI-Native DevOps, this workflow is obsolete. An AI agent has already intercepted the Prometheus alert, analyzed the stack trace, cross-referenced it with recent git commits, and even applied a rollback—all before you woke up. This guide shows you how to build this system.

Architecture: The Loop

The core of AI-Native DevOps is the OODA Loop (Observe, Orient, Decide, Act) automated by an LLM Agent.

The Components

  • Sensor: Prometheus Alertmanager sends a webhook when a metric spikes.
  • Brain: An LLM Agent (LangChain/AutoGPT) receives the webhook.
  • Hands: The Agent has tools: `kubectl`, `git`, `aws-cli`.
  • Judge: A Human-in-the-Loop (via Slack) approves high-risk actions.

Security: The Rogue Agent Problem

Giving an AI `kubectl admin` access is a recipe for disaster. If the LLM hallucinates, it could delete your production database.

Mitigation Strategies

Least Privilege (RBAC)

The Agent's ServiceAccount should only have permission to `get logs` and `restart deployment`. It should never have `delete persistentvolume`.

Human Approval

For any write operation (Act), the agent must send a message to Slack: "I want to restart service X. Approve?"

Building the Agent (Node.js)

This is a simplified webhook handler that acts as our AI DevOps Engineer.

agent.ts (Node.js)
import express from 'express';
import { ChatOpenAI } from "@langchain/openai";
import { initializeAgentExecutorWithOptions } from "langchain/agents";
import { Tool } from "langchain/tools";
import { exec } from 'child_process';

const app = express();
app.use(express.json());

// Define a Custom Tool for Kubectl
class KubectlLogsTool extends Tool {
  name = "kubectl_logs";
  description = "Get logs for a pod. Input should be pod name.";
  async _call(podName: string) {
    return new Promise((resolve) => {
        exec(`kubectl logs ${podName} --tail=50`, (error, stdout) => {
            resolve(stdout || "No logs found or error.");
        });
    });
  }
}

app.post('/webhook/alertmanager', async (req, res) => {
  const alert = req.body.alerts[0];
  const podName = alert.labels.pod;

  const model = new ChatOpenAI({ temperature: 0 });
  const tools = [new KubectlLogsTool()];

  const executor = await initializeAgentExecutorWithOptions(tools, model, {
    agentType: "openai-functions",
  });

  const prompt = `
    Alert received for pod ${podName}.
    The error is: ${alert.annotations.description}.
    1. Get the logs.
    2. Analyze the root cause.
    3. Suggest a fix.
  `;

  const result = await executor.run(prompt);

  // Send result to Slack (omitted)
  console.log("AI Investigation Result:", result);

  res.sendStatus(200);
});

Observability Hooks (Spring Boot)

For the AI to understand the application, the app must expose deep diagnostics. Standard Actuator health checks aren't enough. We need a "Diagnostic Endpoint" designed for LLMs.

DiagnosticController.java
@RestController
public class DiagnosticController {

    @Autowired
    private DataSource dataSource;

    @GetMapping("/internal/diagnose")
    public Map<String, Object> getDiagnostics() {
        Map<String, Object> status = new HashMap<>();

        // 1. DB Connection Pool Status
        if (dataSource instanceof HikariDataSource) {
            HikariPoolMXBean pool = ((HikariDataSource) dataSource).getHikariPoolMXBean();
            status.put("db_active_connections", pool.getActiveConnections());
            status.put("db_threads_awaiting", pool.getThreadsAwaitingConnection());
        }

        // 2. Memory Usage
        status.put("heap_used", Runtime.getRuntime().totalMemory() - Runtime.getRuntime().freeMemory());

        // 3. Last 5 Exceptions (from a memory buffer)
        status.put("recent_errors", ErrorBuffer.getLatest());

        return status;
    }
}

Interactive Incident Workflow

1

Alert Fired

Prometheus detects high latency.

2

Agent Investigates

Agent queries `/internal/diagnose` and sees "db_threads_awaiting: 50".

3

Hypothesis

Agent hypothesizes: "Connection pool exhaustion due to slow query."

4

Remediation

Agent scales up the ReplicaSet to handle load OR restarts the pod (if configured).

Conclusion

AI-Native DevOps is not about replacing SREs; it's about giving them superpowers. By offloading the "investigation" phase to an AI, humans can focus on the "architecture" phase, building systems that are resilient by design.

🌌
Purple Dream
Active Theme