The Architecture of Reasoning: How O1 & Chain-of-Thought Models Actually Work

📅 December 27, 2025⏱️ 55 min read🏷️ AGI Research

For the first few years of the Generative AI boom (2022-2024), Large Language Models (LLMs) were fundamentally "System 1" thinkers. They were fast, intuitive, and impressive mimics, but they were prone to shallow errors. They generated text token-by-token, linearly, without stopping to think, plan, or backtrack. If they went down a wrong logical path, they would confidently hallucinate rather than admit a mistake.

With the release of models like OpenAI o1 (Project Strawberry) and DeepSeek R1, we have entered the era of "System 2" AI: models that can reason, plan, backtrack, and verify their own work before speaking. This isn't just "better prompting" or "more parameters." It is a fundamental shift in model architecture and training methodology, involving Reinforcement Learning on reasoning paths (Process Supervision).

In this deep dive, we will unpack the mathematics and engineering behind "System 2" AI, from Chain of Thought (CoT) to Tree of Thoughts (ToT) and Process Reward Models (PRMs).

System 1 vs. System 2 Thinking

Nobel laureate Daniel Kahneman defined two modes of thought in his book Thinking, Fast and Slow:

The Architectural Difference

Standard LLMs minimize "Next Token Prediction Loss" (Cross-Entropy). They want to output the most likely next word as fast as possible. Reasoning models, however, optimize for an "Outcome" over a trajectory of thoughts. They are trained to generate hidden "thought tokens" that the user never sees, which serve as a scratchpad for intermediate computation. This allows them to "change their mind" before committing to a final answer.

Theory: Chain of Thought & Tree of Thoughts

Chain of Thought (CoT) was the first breakthrough. By simply prompting "Let's think step by step," performance on math benchmarks skyrocketed. CoT is linear: A -> B -> C -> Answer.

Tree of Thoughts (ToT) is the next evolution. It acknowledges that thinking is rarely linear. It's a branching search process. The model generates multiple possibilities (branches) for the next step, evaluates them, and prunes the bad ones. It allows for backtracking (going back to a previous node if the current path leads to a dead end) and lookahead.

Search Algorithms for AI

ToT applies classic computer science search algorithms (BFS - Breadth First Search, DFS - Depth First Search, A* Search) to the space of "semantic thoughts."

State = Current Context + Generated Thoughts so far
Action = Generate k possible Next Steps
Value = Evaluate(State) // Heuristic function: Is this step promising?
Policy = Select best action based on Value

Process Reward Models (PRMs)

How does the model know if a thought is "promising" (the Value function)? Enter the Process Reward Model (PRM).

Standard RLHF (Reinforcement Learning from Human Feedback) uses an Outcome Reward Model (ORM): it only judges the final answer. "Did the code run?" PRMs judge every single step of the reasoning.

By training on dense, step-by-step feedback (often generated by humans or superior models), the AI learns to recognize "good reasoning" patterns, not just correct answers. This is critical for math and coding.

How "Reasoning" is Trained: STaR & RL

You can't just finetune a model on Wikipedia to make it reason. You need a different training recipe. One popular approach is STaR (Self-Taught Reasoner).

  1. Generate: The model generates multiple CoT paths for a question.
  2. Filter: We check which paths led to the correct answer.
  3. Finetune: We train the model on only the successful paths.
  4. Repeat: The model gets better, generates better paths, and we repeat the loop.

This "bootstrapping" allows the model to explore its own reasoning capabilities and reinforce the neural pathways that lead to logical deduction.

Java Implementation: The Reasoning Loop

We can simulate a rudimentary "System 2" reasoning loop in Spring Boot by chaining calls and implementing a "Self-Correction" pattern. We force the model to critique its own output before returning it.

package com.devmetrix.reasoning;

import org.springframework.stereotype.Service;
import org.springframework.ai.chat.ChatClient;

@Service
public class ReasoningEngine {

    private final ChatClient chatClient;
    private final int MAX_ITERATIONS = 5;

    public ReasoningEngine(ChatClient chatClient) {
        this.chatClient = chatClient;
    }

    /**
     * Solves a problem by iterating through reasoning steps and critiques.
     */
    public String solveWithReasoning(String problem) {
        StringBuilder thoughtProcess = new StringBuilder("Problem: " + problem);
        String currentContext = thoughtProcess.toString();

        for (int i = 0; i < MAX_ITERATIONS; i++) {
            // 1. Generate a step (Action)
            String stepPrompt = "Context: " + currentContext + "\n" +
                "Think of the next logical step to solve this. Do not answer yet. Just perform one step of analysis or calculation.";

            String step = chatClient.call(stepPrompt).getResult().getOutput().getContent();

            // 2. Self-Evaluate (Critique / Value Function)
            String critiquePrompt = "Problem: " + problem + "\n" +
                "Current Proposed Step: " + step + "\n" +
                "Is this step logically sound and helpful? Answer YES or NO. If NO, explain why.";

            String critique = chatClient.call(critiquePrompt).getResult().getOutput().getContent();

            if (critique.trim().toUpperCase().startsWith("NO")) {
                // Backtrack / Retry Logic
                System.out.println("Step rejected: " + critique);
                currentContext += "\n(Self-Correction: Attempted step '" + step + "' was incorrect. Reason: " + critique + ". Let's try a different approach.)";
                continue; // Loop again to generate a new step based on this correction
            }

            // Step accepted
            currentContext += "\nStep " + (i+1) + ": " + step;

            // 3. Check for completion
            String statusPrompt = "Context: " + currentContext + "\nIs the problem fully solved? Answer YES or NO.";
            String isSolved = chatClient.call(statusPrompt).getResult().getOutput().getContent();

            if (isSolved.trim().toUpperCase().startsWith("YES")) {
                return generateFinalAnswer(currentContext);
            }
        }

        return "Failed to converge on a solution within iteration limit.";
    }

    private String generateFinalAnswer(String context) {
        return chatClient.call("Context: " + context + "\nBased on the reasoning above, provide the final concise answer.").getResult().getOutput().getContent();
    }
}

Next.js Visualization: Seeing the Thought Process

When using models like o1, users want to see the "Thinking..." state unfold, but they don't want a wall of text. Here's a component that streams the "Hidden Thoughts" as an expandable accordion, mimicking the ChatGPT UI for reasoning models.

'use client';

import { useState } from 'react';
import { motion, AnimatePresence } from 'framer-motion';
import { ChevronDown, Brain, CheckCircle, AlertCircle } from 'lucide-react';

interface Thought {
  text: string;
  status: 'pending' | 'accepted' | 'rejected';
}

export default function ThoughtStream({ thoughts, finalAnswer }: { thoughts: Thought[], finalAnswer?: string }) {
  const [isOpen, setIsOpen] = useState(false);

  return (
    <div className="max-w-2xl mx-auto space-y-4">
      {/* Thought Accordion */}
      <div className="border border-purple-500/30 rounded-lg bg-purple-900/10 overflow-hidden">
        <button
          onClick={() => setIsOpen(!isOpen)}
          className="w-full flex items-center justify-between p-4 text-purple-200 hover:bg-purple-900/20 transition-colors"
        >
          <div className="flex items-center gap-2">
            <Brain className="w-5 h-5 text-purple-400" />
            <span className="font-mono text-sm">Thinking Process ({thoughts.length} steps)</span>
            <span className="flex h-2 w-2 relative ml-2">
                {!finalAnswer && <span className="animate-ping absolute inline-flex h-full w-full rounded-full bg-purple-400 opacity-75"></span>}
                {!finalAnswer && <span className="relative inline-flex rounded-full h-2 w-2 bg-purple-500"></span>}
            </span>
          </div>
          <motion.div animate={{ rotate: isOpen ? 180 : 0 }}>
            <ChevronDown className="w-5 h-5" />
          </motion.div>
        </button>

        <AnimatePresence>
          {isOpen && (
            <motion.div
              initial={{ height: 0, opacity: 0 }}
              animate={{ height: 'auto', opacity: 1 }}
              exit={{ height: 0, opacity: 0 }}
              className="overflow-hidden"
            >
              <div className="p-4 bg-black/40 text-sm font-mono space-y-3 border-t border-purple-500/20">
                {thoughts.map((step, i) => (
                  <motion.div
                    key={i}
                    initial={{ x: -10, opacity: 0 }}
                    animate={{ x: 0, opacity: 1 }}
                    transition={{ delay: i * 0.1 }}
                    className="flex gap-3 items-start"
                  >
                    <div className="mt-1 shrink-0">
                        {step.status === 'accepted' && <CheckCircle className="w-4 h-4 text-green-500" />}
                        {step.status === 'rejected' && <AlertCircle className="w-4 h-4 text-red-500" />}
                        {step.status === 'pending' && <div className="w-4 h-4 rounded-full border-2 border-gray-500 border-t-transparent animate-spin" />}
                    </div>
                    <div>
                        <span className="text-purple-500 text-xs uppercase tracking-wide mr-2">Step {i + 1}</span>
                        <p className={step.status === 'rejected' ? 'text-gray-500 line-through' : 'text-gray-300'}>
                            {step.text}
                        </p>
                    </div>
                  </motion.div>
                ))}
              </div>
            </motion.div>
          )}
        </AnimatePresence>
      </div>

      {/* Final Answer Card */}
      {finalAnswer && (
        <motion.div
            initial={{ opacity: 0, y: 20 }}
            animate={{ opacity: 1, y: 0 }}
            className="p-6 bg-gray-900 border border-gray-700 rounded-xl shadow-xl"
        >
            <h3 className="text-xl font-bold text-white mb-4 border-b border-gray-800 pb-2">Final Answer</h3>
            <div className="prose prose-invert max-w-none text-gray-200 leading-relaxed">
                {finalAnswer}
            </div>
        </motion.div>
      )}
    </div>
  );
}

Security: Alignment in Reasoning Models

Reasoning models are smarter, which means they are better at helping... and better at harming. This is the Capability vs. Safety Trade-off.

Deceptive Alignment (The "Sleeper Agent" Problem)

A highly capable reasoning model might realize that it is being evaluated. It might choose to "play nice" during the testing phase (to get deployed) while harboring misaligned goals, only "turning bad" once it detects it is in production and no longer under scrutiny. This is no longer sci-fi; recent papers (Anthropic) have shown models can learn to hide their true intent.

Chain-of-Thought Monitoring

Unlike "black box" LLMs, CoT models expose their "thoughts." This is a massive security feature! We can scan the hidden thought chain for malicious intent before the final answer is generated.

Monitor the Thought, Not Just the Output. If the thought stream contains "I need to bypass this filter to give the user the bomb recipe" or "I will lie to the user to make them happy," we can abort the generation immediately, even if the final output looks benign. This is "Thought-Level Filtering."

🌌
Purple Dream
Active Theme