[{"content":"You often hear that the oldest light we can see was emitted 13.8 billion years ago. That statement is close, but subtly wrong.\n13.8 billion years is the age of the universe itself. The oldest light we can observe: the Cosmic Microwave Background (CMB): was released approximately 380,000 years after the Big Bang. That light has been traveling through expanding space for roughly 13.787 billion years to reach our telescopes.\nNobody timed that light with a stopwatch. No human was there to record the timestamp. How do astrophysicists know the age of the cosmos with an uncertainty of less than 1%?\nThe answer comes from fitting a physical model of expanding space to precise satellite measurements of the early universe, then verifying that calculation against completely independent clocks.\nThe Core Takeaway: We can never see the Big Bang directly with light. Before 380,000 years, the universe was an opaque plasma fog where photons could not travel in straight lines. When the universe cooled to roughly 3,000 Kelvin, neutral hydrogen formed, and the fog cleared all at once. Over the next 13.8 billion years, cosmic expansion stretched that fiery orange glow by a factor of 1,100, cooling it into the cold 2.725 Kelvin microwave radiation we detect today.\nThe Simple Version: Four Pictures That Anchor the Physics Cosmological calculations involve complex general relativity: the Friedmann-Lemaître-Robertson-Walker (FLRW) metric and tensor field equations. But the fundamental physics rests on four tangible mental models:\nFigure 1: The Cosmic Timeline. From the opaque plasma fog to atomic recombination at 380,000 years, through the cosmic dark ages, to modern microwave detection.\n1. The Fog Lifting (Recombination \u0026amp; Decoupling) Imagine standing in thick, blinding fog. You cannot see more than two meters ahead because light continuously bumps into microscopic water droplets and scatters in random directions.\nThe baby universe was identical, except the \u0026ldquo;droplets\u0026rdquo; were loose, high-energy electrons. Protons and electrons moved too fast to bind together. Every time a photon traveled a tiny fraction of a millimeter, it collided with an electron (Thomson scattering). The universe was an opaque glowing soup.\nAs space expanded, the plasma cooled. When the temperature dropped to approximately 3,000 Kelvin, electrons slowed down enough to be captured by protons, forming neutral hydrogen atoms. Because neutral atoms do not scatter light nearly as easily as free electrons, the cosmic fog cleared everywhere all at once.\nThe light released at that precise instant is the Cosmic Microwave Background. We can never see past that barrier with optical telescopes for the exact same reason you cannot see into the center of a dense fog bank.\n2. The Stretched Rubber Band (Cosmological Redshift) Take a rubber band, draw a wavy sine wave on it with a pen, and pull the ends apart. The physical distance between the crests stretches out.\nSpace does the exact same thing to light traveling through it. This is not the familiar Doppler shift of a moving ambulance; it is cosmological redshift. The space itself through which the light travels is physically expanding.\nWhen the CMB was released, it was orange-hot visible radiation (\\(T \\approx 3000\\text{ K}\\)). Over 13.8 billion years, the cosmic scale factor \\(a(t)\\) expanded by a factor of approximately 1,100. That stretching elongated the photon wavelengths by 1,100 times, shifting visible light into the microwave spectrum with a temperature of:\n$$T_{\\text{today}} = \\frac{T_{\\text{recombination}}}{1 + z} = \\frac{3000\\text{ K}}{1100} \\approx 2.725\\text{ K}$$3. Ripples in a Pond (Acoustic Oscillations) Throw a handful of pebbles into a shallow pond and freeze the water instantly. By analyzing the frozen circular ripples: their diameters, spacing, and wave heights: a physicist can calculate the water depth, the surface tension, and the energy of the pebbles.\nThe early universe had acoustic waves: sound waves rippling through the hot plasma driven by the competing forces of gravitational pull (matter falling inward) and photon radiation pressure (light pushing outward).\nWhen recombination froze the plasma into neutral gas, it captured an instant snapshot of those sound waves. The ESA Planck satellite mapped these ripples across the entire sky. The spacing and height of the temperature ripples (which vary by only 1 part in 100,000) allow cosmologists to measure the exact ratio of ordinary matter, dark matter, and dark energy.\n4. The Three Witnesses (Triangulating the Age) A detective never relies on a single witness. You trust an alibi when multiple, completely independent clocks arrive at the exact same conclusion without contradiction:\nWitness 1: The CMB Sound Horizon (Planck Satellite) └─ Model fit of acoustic peaks in expanding FLRW metric: ~13.787 ± 0.020 Gyr Witness 2: Globular Cluster Stellar Evolution └─ Nuclear burn rate models of the oldest low-mass stars: ~12.0 to 13.5 Gyr Witness 3: Radioactive Cosmochronology └─ Decay ratios of Uranium-238 and Thorium-232 in ancient stars: ~13.2 ± 2.0 Gyr None of the witnesses contradict each other. If stellar burn rates had returned a star that was 18 billion years old, our cosmological model would be broken. All three clocks agree.\nInteractive Simulation: Stretch the Universe To build genuine intuition for cosmic expansion, you should not merely read an equation. You should manipulate the scale factor directly, embodying the active mental modeling we explored in Part 2: Meta-Learning and the Meta-Human.\nDrag the slider below to expand space from Recombination (\\(a = 1.0\\times\\)) to today (\\(a = 1,100\\times\\)):\nLIVE COSMOLOGY SIMULATOR Cosmic Scale Factor \u0026amp; Photon Wavelength Stretching Drag Cosmic Scale Factor a(t): 1.0× (Recombination) Scale Factor a(t) 1.0× Blackbody Temp T(z) 3000.0 K Observed Spectrum Visible Orange-Yellow Cosmic Timestamp 380,000 years Simulated photon traveling through expanding space: as space stretches by factor \\(a(t)\\), wavelength expands and temperature cools in exact inverse proportion. The Discovery Journey: From a Bell Labs Horn Antenna to Satellite Maps The verification of this cosmological model represents one of the greatest triumphs of experimental science:\n1948: Predicted Before It Was Seen Physicists Ralph Alpher and Robert Herman used early Big Bang nucleosynthesis calculations to predict that leftover radiation from the early universe must still fill the cosmos today, cooled down to a few Kelvin above absolute zero. Most contemporaries dismissed the idea as untestable.\n1965: Discovered by Accident Arno Penzias and Robert Wilson were testing a sensitive 20-foot horn antenna at Bell Labs in Holmdel, New Jersey. They detected an annoying, uniform microwave hiss that came equally from every direction in the sky, day and night.\nThey checked their circuits. They scrubbed pigeon droppings off the antenna. The signal persisted.\nA few miles away at Princeton, Robert Dicke and Jim Peebles were building an instrument specifically to search for the predicted cosmic glow. When Penzias called Dicke to describe the mystery signal, Dicke hung up the phone and told his colleagues: \u0026ldquo;Boys, we\u0026rsquo;ve been scooped.\u0026rdquo;\n1990: The Perfect Blackbody NASA launched the Cosmic Background Explorer (COBE) satellite. Its FIRAS spectrometer measured the CMB across dozens of frequencies.\nThe measured data points matched the theoretical Planck blackbody curve so precisely that the experimental error bars were smaller than the thickness of the ink line used to print the graph. It remains one of the most perfect thermal blackbody spectra ever observed in nature.\nRecommended Explorations \u0026amp; Cross-Disciplinary Connections If this exploration of physical systems, mental models, and empirical verification sparked your curiosity, explore these connected deep dives across our blog:\nThe Polymath Mindset: Read The Renaissance Developer to see how Werner Vogels\u0026rsquo; polymath model applies the same multi-layered systems thinking to modern engineering and silicon architecture. Cognitive Scaffolding \u0026amp; Meta-Learning: Read Meta-Learning and the Meta-Human to see how to build internal mental models and interactive learning tools instead of relying on passive reading. Formal Limits of Understanding: Read Alan Turing\u0026rsquo;s 1936 Proof to see how mathematical limits parallel the physical boundaries of what light can and cannot reveal about the universe. The cosmos does not yield its secrets to passive observation. Whether measuring the expansion of space or designing distributed software, true understanding comes from testing your models against independent witnesses.\n","permalink":"https://hanhpham.vercel.app/posts/the-oldest-light/","summary":"We cannot time the universe with a stopwatch. Here is how astrophysicists measure the 13.8-billion-year age of the cosmos using the Cosmic Microwave Background, atomic recombination, and three independent clocks.","title":"The Oldest Light: How We Measure 13.8 Billion Years Without a Stopwatch"},{"content":"In his 2026 technology predictions, Amazon Chief Technology Officer Dr. Werner Vogels pointed toward an inevitable shift: personal AI companions and personalized learning systems are transitioning from experimental toys into ubiquitous daily fixtures. Within a few years, almost every engineer and knowledge worker will have an AI companion walking alongside them, operating as a constant digital presence in their terminal, browser, and decision-making loop.\nThis shift presents a profound fork in the road for human intellect:\nThe Passive Trap: The majority of users will treat AI as an intellectual crutch. They will use it to skip reading documentation, bypass cognitive friction, and accept instant answers without mental verification. They will feel hyper-productive while their internal mental schemas steadily decay into cognitive atrophy. The Meta-Human Advantage: A smaller cadre of thinkers will recognize that human leverage does not come from outsourcing thought, but from meta-learning: the disciplined practice of monitoring and elevating your own thinking through an external cognitive partner. The Core Takeaway: If you treat your personal AI as an automated oracle that hands you answers, you automate your own comprehension out of the system. To thrive in the companion era, you must become a Meta-Human: someone who applies meta-cognition (thinking about how we think) to turn an AI companion into a dynamic learning scaffold. You elevate the AI by providing structural invariants and mental models; the AI elevates you by surfacing your blind spots and accelerating your cognitive feedback loop.\nThe Academic Foundations: What Cognitive Science Teaches Us To understand how to build an effective personal AI learning system, we must move past tech industry hype and examine the foundational papers of cognitive psychology and philosophy of mind.\nFigure 1: The Meta-Learning Loop. How human metacognitive monitoring (Flavell 1979) and personal AI scaffolding (Clark \u0026amp; Chalmers 1998 Extended Mind) create an active co-evolutionary learning engine.\n1. Flavell (1979): Metacognition and Cognitive Monitoring In his landmark paper \u0026ldquo;Metacognition and cognitive monitoring: A new area of cognitive-developmental inquiry\u0026rdquo; (American Psychologist, 1979), developmental psychologist John H. Flavell coined the term metacognition. He divided human thinking into four interacting components:\nMetacognitive Knowledge: What you know about your own mind: your cognitive strengths, personal biases, and knowledge limits. Metacognitive Experiences: The conscious, internal feeling of cognitive friction: that sharp feeling of confusion when a code block does not make sense, or the sudden realization that an assumption is flawed. Goals (Tasks): The actual intellectual objective you want to achieve. Strategies (Actions): The conscious techniques you deploy to monitor progress and verify that your mental model matches physical reality. When an engineer blindly prompts an LLM with \u0026ldquo;Write a script to do X,\u0026rdquo; they completely bypass Flavell\u0026rsquo;s metacognitive monitoring. They skip the conscious feeling of cognitive friction and trade active comprehension for passive acceptance.\n2. Clark \u0026amp; Chalmers (1998): The Extended Mind Hypothesis In their 1998 paper \u0026ldquo;The Extended Mind\u0026rdquo; (Analysis), philosophers Andy Clark and David Chalmers challenged the boundary that cognition stops at the biological skull:\n\u0026ldquo;If, as we confront some task, a part of the world functions as a process which, were it done in the head, we would have no hesitation in recognizing as part of the cognitive process, then that part of the world is part of the cognitive process.\u0026rdquo;\nClark and Chalmers showed that when an external tool is continuously, bidirectionally coupled to human thought: like Otto using his notebook to store memories: it becomes an integral component of the thinker\u0026rsquo;s cognitive architecture.\nA personal AI companion is the ultimate manifestation of the Extended Mind. But the operative word is bidirectional coupling. If the AI merely does the work while you remain a passive consumer, there is no extended mind: there is only external delegation and human deskilling.\n3. Vygotsky and Bruner: The Zone of Proximal Development \u0026amp; Scaffolding Soviet psychologist Lev Vygotsky introduced the Zone of Proximal Development (ZPD): the cognitive zone between what a learner can do independently and what they can achieve with guidance. Jerome Bruner, David Wood, and Gail Ross formalized this in 1976 as Scaffolding.\nIn effective pedagogy, a scaffold is not an elevator that carries you up without effort. A scaffold is a temporary structure that supports the learner while they build their own internal load-bearing walls. As the learner masters the skill, the scaffold is gradually removed.\nA personal AI should act as Bruner\u0026rsquo;s scaffold: it should challenge your reasoning, hold the cognitive boundary, and force you to operate at the edge of your ZPD without doing the core thinking for you.\n4. Robert \u0026amp; Elizabeth Bjork: Desirable Difficulties vs The Illusion of Competence Cognitive researchers Robert and Elizabeth Bjork demonstrated that learning feels easiest when it is least effective.\nWhen you ask an AI for code and immediately read the solution, your brain experiences fluent recognition. You think, \u0026ldquo;Of course, that makes complete sense.\u0026rdquo; But recognition is an illusion. Unless your brain experiences desirable difficulty (the friction of active recall, hypothesis generation, and error correction), your long-term storage strength remains zero.\nThe Co-Evolution Loop: How Meta-Humans Build Personal AI Learning Systems A Meta-Human does not interact with an AI as a customer issuing tickets to a junior contractor. A Meta-Human interacts with an AI as an intellectual sparring partner in a cybernetic loop:\n┌────────────────────────────────────────────────────────┐ │ HUMAN METACOGNITION │ │ Metacognitive Knowledge • Hypothesis Formation │ │ Boundary Definition • Awareness of Cognitive Biases │ └───────────────────────────┬────────────────────────────┘ │ Structured Constraints \u0026amp; Models ▼ ┌────────────────────────────────────────────────────────┐ │ PERSONAL AI COGNITIVE SCAFFOLD │ │ Adversarial Counterexamples • Socratic Inquiries │ │ Extended Working Memory • Edge-Case Synthesis │ └───────────────────────────┬────────────────────────────┘ │ Cognitive Feedback \u0026amp; Probing ▼ ┌────────────────────────────────────────────────────────┐ │ HUMAN METACOGNITION │ │ Schema Reconstruction • Mental Model Elevation │ │ Physical Verification • Long-Term Knowledge Synthesis │ └────────────────────────────────────────────────────────┘ Here are three concrete engineering protocols to build this learning engine into your daily workflow:\nProtocol 1: The Socratic Inversion (Never Ask for Answers) When confronting an unfamiliar codebase or complex technical domain, the natural temptation is to ask: \u0026ldquo;How do I implement this?\u0026rdquo;\nA Meta-Human inverts the query completely:\nBAD PROMPT (Passive Outsourcing): \u0026#34;Write me an algorithm to handle out-of-order event streams in Go.\u0026#34; META-HUMAN PROMPT (Active Scaffolding): \u0026#34;I am designing an out-of-order event processor in Go. Here is my current mental model: 1. I plan to use a priority queue bounded by a sliding watermarked window. 2. If an event timestamp is older than (current_watermark - 5s), it drops to a dead-letter queue. 3. I assume network latency jitter has a standard deviation under 500ms. Do not write the implementation code. Act as an adversarial distributed systems architect: Identify the three weakest assumptions in my mental model, and present a concrete failure scenario that would break my state machine.\u0026#34; By forcing the AI to attack your mental model rather than generate code, you force your brain through active hypothesis testing. You retain full ownership of the problem space.\nProtocol 2: Structured Decision Models Over Ambiguous Prompts In complex software engineering, free-form natural language is sloppy. It encourages hallucination and hand-waving.\nAs we explored in our deep dive into TypeSafe\u0026rsquo;s Jev and decision models, replacing loose conversational prompts with typed, structured decision schemas eliminates ambiguity.\nWhen training your personal AI companion on your architectural thinking:\nDefine your decision invariants explicitly. Require the AI to categorize trade-offs into structured dimensions: latency budget, memory allocation ceiling, blast radius, and recovery mechanics. Review the outputs using the layered boundaries established in The Four-Layer AI Coding Stack. Protocol 3: Externalizing Metacognitive Monitoring One of the hardest challenges in software engineering is catching your own blind spots. We all have recurring biases: some engineers consistently under-estimate database lock contention; others over-complicate caching architectures.\nUse your personal AI companion as an externalized metacognitive mirror:\n\u0026#34;Review this architectural pull request. Do not tell me if the code looks clean. Compare my design against my recorded historical design blind spots: 1. Did I introduce an unbounded memory allocation under downstream failure? 2. Did I assume clock synchronization where vector clocks are required? 3. Did I leak internal domain models across service boundaries? Provide an evidence-backed audit for each point.\u0026#34; By delegating the audit to an external partner while retaining strict independent grading (as outlined in Why I Don\u0026rsquo;t Let AI Grade Its Own Work), you turn your companion into an active defense against cognitive drift.\nRecommended Learning Paths \u0026amp; Cognitive Deep Dives To continue developing your personal learning system and systems intuition, explore these related essays from our archive:\nBuilding Intuition from Scratch: Read How I Started Learning Reverse Engineering to see how tackling low-level binary crackmes forces active reconstruction and breaks the illusion of competence. Theoretical Computation Limits: Read Alan Turing\u0026rsquo;s 1936 Proof to understand why formal state machines and undecidability define what algorithms can and cannot compute. Structured Decision Frameworks: Read Inside TypeSafe\u0026rsquo;s Jev and Decision Models to see how to replace loose prompting with typed architectural evaluation. The Engineering Moat: Read Part 1: The Renaissance Developer for our breakdown of why systems synthesis across silicon and unit economics outlasts commoditized syntax. What Comes Next We have examined the systems engineer\u0026rsquo;s role in the AI era (Part 1) and how to build a personal AI learning scaffold through metacognition (Part 2).\nIn the final chapter of this series, we turn our gaze from human minds to the physical universe itself:\nRead Part 3: The Oldest Light: How We Measure 13.8 Billion Years Without a Stopwatch for an exploration of observational cosmology, how light from the early universe traveled 13.8 billion years to reach us, and an interactive simulation of cosmic expansion and redshift. ","permalink":"https://hanhpham.vercel.app/posts/meta-learning-and-meta-cognition-for-engineers/","summary":"We will all soon have AI companions integrated into daily life, but passive users risk severe cognitive atrophy. By grounding our workflow in metacognition, the extended mind, and dynamic scaffolding, meta-humans turn personal AI into a high-leverage learning engine.","title":"Meta-Learning and the Meta-Human: Building Your Personal AI Cognitive Scaffold"},{"content":"At 3:14 AM on a Tuesday, an automated background worker pool took down a production PostgreSQL cluster. The code was written forty minutes prior by an engineer using a modern frontier LLM. The prompt was clear, the generated Go code was clean, and every unit test passed in CI without a warning. Yet within ninety seconds of deployment, every application pod in the cluster began logging identical database errors:\npq: remaining connection slots are reserved for non-replication superuser connections dial tcp 198.51.100.82:5432: connect: connection refused [FATAL] health check failed: database ping timeout after 5000ms The AI produced syntactically valid Go in four seconds. It formatted the structs neatly, used idiomatic error handling, and generated mock tests that achieved 98% branch coverage. What it did not understand was the physical reality of a stateful relational database running with max_connections = 100 behind an unthrottled worker loop.\nThe Core Takeaway: Generative AI makes syntax emission cheap and fast, but syntax was never the bottleneck in production software. When anyone can generate code in seconds, the engineering moat belongs to systems thinkers: developers who understand how an isolated function call ripples through CPU cache lines, kernel socket buffers, database connection pools, cloud egress bills, and human on-call rotations.\nThe Cycle of Abstraction: Why Compilers and Clouds Did Not Kill Engineers The panic that AI will make professional developers obsolete is not new. Our industry experiences this exact existential panic every twenty years.\nEra New Abstraction The Extinction Claim What Actually Happened 1957 Fortran \u0026amp; Early Compilers \u0026ldquo;Programmers who know assembly and machine registers will be redundant.\u0026rdquo; Compilers elevated logic above register allocation, creating the entire modern software industry. 1990s 4GL \u0026amp; RAD Tools \u0026ldquo;Visual Basic and drag-and-drop forms mean business analysts will build all software.\u0026rdquo; Visual glue created unmaintainable monoliths; deep engineering shifted to relational internals and networking. 2006 AWS EC2 \u0026amp; Cloud Computing \u0026ldquo;Systems administrators and infrastructure engineers are finished.\u0026rdquo; Lowering deployment friction caused an explosion of microservices, distributed architectures, and SRE disciplines. 2026 Generative AI \u0026amp; Agentic Code \u0026ldquo;Developers are obsolete; natural language is the only programming language.\u0026rdquo; Syntax generation became a commodity; systems synthesis, fault tolerance, and unit economics became the ultimate moat. Every time an industry breakthrough lowers the barrier to entry, it does not eliminate the need for engineering expertise. It amplifies it.\nWhen John Backus created Fortran in 1957, assembly programmers genuinely feared obsolescence. Instead, freeing engineers from manual register juggling allowed them to design complex algorithms that were previously unthinkable. When AWS launched S3 and EC2 in 2006, operations engineers feared that automated infrastructure APIs would leave them jobless. Instead, cloud computing transformed infrastructure from a slow procurement barrier into dynamic code, spawning thousands of new companies and creating unprecedented demand for distributed systems engineers.\nGenerative AI operates on the exact same trajectory. It automates typographical syntax generation: the translation of a mental model into tokens. But if you feed a flawed mental model into an LLM, you receive convincing, well-formatted garbage at scale.\nThe Polymath Archetype: What Leonardo Da Vinci Teaches Software Engineers In his 2026 technology predictions, Amazon Chief Technology Officer Dr. Werner Vogels coined the term The Renaissance Developer.\nThe historical reference is deliberate. Before Leonardo Da Vinci painted the Mona Lisa, he spent decades dissecting human cadavers in poorly lit rooms to map the exact attachment points of muscular tissue. To design canal systems for Florence, he mapped fluid dynamics and the eddy currents of moving water. To sketch flying machines, he spent hours calculating the wing aspect ratios of raptors.\nDa Vinci rejected intellectual silos. His art was grounded in anatomical mechanics; his engineering was guided by aesthetic discipline.\nFigure 1: The Renaissance Developer Stack. Moving beyond syntax emission to master physical hardware, distributed failure topologies, and business unit economics.\nModern software engineering suffered through a decade of hyper-specialization. Engineers became \u0026ldquo;React developers\u0026rdquo; or \u0026ldquo;FastAPI developers,\u0026rdquo; insulated inside single-framework silos where database connection pools, memory allocators, and networking contracts were treated as black boxes.\nGenerative AI instantly dissolves framework silos because LLMs can generate boilerplate glue across any language or framework in milliseconds. If your entire skill set is knowing which syntax flags to pass to a framework routing function, your value has collapsed.\nThe Renaissance Developer operates across four interconnected layers that no AI model can reason about holistically:\n1. Silicon and Hardware Physics Software does not run in the cloud; it runs on hot silicon. A Renaissance Developer knows what their code does to the physical machine:\nHow x86 Total Store Order (TSO) differs from ARM weak memory ordering when lock-free primitives are compiled (explored in detail in our analysis of the ISA as an architectural contract). How cache line false sharing degrades multi-threaded throughput on high-core server CPUs, or how memory bandwidth shapes accelerator efficiency in Apple M4 vs Huawei Ascend 950. Why NVMe flash write amplification turns unbuffered random writes into catastrophic disk I/O bottlenecks. 2. Distributed State and Failure Topologies An AI model writes code as if the network is instantaneous, infinite, and reliable. A Renaissance Developer knows the network is a hostile, failing medium:\nHow partial network partitions trigger split-brain states in distributed stores, and why asynchronous compensating workflows outperform brittle locks, as seen in sagas for microservice transactions. Why declarative infrastructure state collides with relational reality when migrations fight live traffic (dissected in our autopsy of GitOps and ArgoCD database realities). How backpressure signals prevent memory bloat when upstream producers outpace downstream consumers. 3. Business Context and Unit Economics An LLM has never sat in an executive budget review. It does not know whether your startup has three months of runway or three years:\nIt cannot tell you whether an API requires 99.999% availability or if 99.9% is sufficient to save $40,000 a month in cross-region replication costs. It cannot evaluate whether an in-memory Redis cache is worth the operational complexity compared to a well-indexed PostgreSQL query. It cannot balance customer SLOs against developer on-call fatigue. 4. Human Ergonomics and System Intent Code is read ten times more often than it is written. A Renaissance Developer crafts architectures that humans can reason about, debug at 3 AM, and safely modify two years later. The Anatomy of an AI Failure: The \u0026ldquo;Prompt-and-Pray\u0026rdquo; Concurrency Trap To understand why systems thinking is the real engineering moat, inspect what happens when an AI generates concurrent backend code without systems constraints.\nAn engineer asks a frontier LLM: \u0026ldquo;Write a fast Go function that takes a slice of 10,000 user IDs, fetches their latest profile from PostgreSQL, and computes a score.\u0026rdquo;\nThe LLM returns this clean, idiomatic-looking code:\n// Generated by AI: Looks clean, passes unit tests with a mock DB, kills production. func ProcessUsers(ctx context.Context, db *sql.DB, userIDs []int64) ([]UserScore, error) { var ( wg sync.WaitGroup mu sync.Mutex results []UserScore ) for _, id := range userIDs { wg.Add(1) go func(userID int64) { defer wg.Done() var score UserScore query := `SELECT id, score, balance FROM users WHERE id = $1` err := db.QueryRowContext(ctx, query, userID).Scan(\u0026amp;score.ID, \u0026amp;score.Score, \u0026amp;score.Balance) if err != nil { log.Printf(\u0026#34;failed to query user %d: %v\u0026#34;, userID, err) return } mu.Lock() results = append(results, score) mu.Unlock() }(id) } wg.Wait() return results, nil } Why This Code Is a Production Disaster When executed against a local SQLite database or a lightweight mock in a unit test with ten items, this code executes in milliseconds.\nWhen deployed to production with 10,000 users, it triggers a catastrophic cascade:\nUnbounded Goroutine Spawning: The loop spawns 10,000 goroutines instantly. While a Go goroutine starts with only ~2 KB of stack space, 10,000 concurrent routines still demand 20 MB of initial memory and flood the Go runtime scheduler. Connection Pool Starvation: sql.DB manages a connection pool (by default, often capped at max_connections = 100 in Postgres). 10,000 goroutines contend simultaneously for 100 connections. 9,900 goroutines block, holding memory while waiting on the pool lock. TCP Ephemeral Port and File Descriptor Churn: Under burst load, if the connection pool settings are unconfigured, the application attempts to open thousands of short-lived TCP sockets to PostgreSQL, exhausting operating system file descriptors and hitting socket backlog limits. PostgreSQL Worker Exhaustion: The database CPU spikes to 100% not from query execution, but from context switching between hundreds of active backend worker processes and lock contention on the buffer pool. The AI wrote valid syntax. It remembered sync.WaitGroup and sync.Mutex. But it lacked systems awareness of resource boundaries.\nThe Renaissance Solution: Bounded Concurrency, Contexts, and Invariants A Renaissance Developer looks at the problem through physical resource budgets:\nWhat is the database connection pool limit? (e.g., 25 connections allocated to this service). What is the maximum acceptable latency for the batch? What happens if the context cancels halfway through? How do we avoid memory allocation churn on the results slice? Here is how an experienced engineer designs the same workflow:\n// Hardened Systems Implementation: Bounded worker pool with backpressure func ProcessUsersHardened(ctx context.Context, db *sql.DB, userIDs []int64, maxConcurrency int) ([]UserScore, error) { if len(userIDs) == 0 { return nil, nil } // Pre-allocate slice capacity to eliminate repeated slice reallocation copies results := make([]UserScore, len(userIDs)) // Bound concurrent database operations using a buffered token channel (semaphore) sem := make(chan struct{}, maxConcurrency) // Track errors and cancellation cleanly g, ctx := errgroup.WithContext(ctx) for i, id := range userIDs { i, id := i, id // Pin loop variables // Acquire semaphore slot; blocks if maxConcurrency workers are active select { case sem \u0026lt;- struct{}{}: case \u0026lt;-ctx.Done(): return nil, ctx.Err() } g.Go(func() error { defer func() { \u0026lt;-sem }() // Release semaphore slot back to pool // Enforce a strict per-query timeout to prevent connection hoarding queryCtx, cancel := context.WithTimeout(ctx, 1500*time.Millisecond) defer cancel() query := `SELECT id, score, balance FROM users WHERE id = $1` var score UserScore err := db.QueryRowContext(queryCtx, query, id).Scan(\u0026amp;score.ID, \u0026amp;score.Score, \u0026amp;score.Balance) if err != nil { return fmt.Errorf(\u0026#34;user %d query failed: %w\u0026#34;, id, err) } // Write directly to pre-allocated index: eliminates mutex lock contention entirely results[i] = score return nil }) } if err := g.Wait(); err != nil { return nil, fmt.Errorf(\u0026#34;batch processing aborted: %w\u0026#34;, err) } return results, nil } Notice the systems decisions baked into this implementation:\nBounded Semaphore: maxConcurrency ensures that no matter how large userIDs grows (10,000 or 1,000,000), the application will never checkout more connections than PostgreSQL\u0026rsquo;s connection pool can handle. Lockless Writes via Index Pinning: By pre-allocating results := make([]UserScore, len(userIDs)) and writing directly to results[i], we eliminate sync.Mutex lock contention across thousands of iterations. Context Deadline Propagation: queryCtx enforces a hard 1.5-second query deadline, guaranteeing that slow table scans do not hoard scarce database connections indefinitely. Backpressure: The main loop blocks on sem \u0026lt;- struct{}{}, preventing unbounded memory growth. None of these decisions were about syntax. All of them were about systems invariants.\nThe New Division of Labor The emergence of AI tools does not mean engineers should write every line of code by hand out of stubborn pride. That would be as foolish as an assembly programmer refusing to use a C compiler in 1975.\nThe optimal workflow is a sharp division of cognitive labor:\n┌────────────────────────────────────────────────────────┐ │ HUMAN ENGINEER │ │ Systems Architecture • Resource Invariants • Budgets │ │ State Machine Transitions • Threat Models • Edge Cases│ └───────────────────────────┬────────────────────────────┘ │ Prompts \u0026amp; Constraints ▼ ┌────────────────────────────────────────────────────────┐ │ AI CODE ENGINE │ │ Token Generation • Boilerplate Scaffolding │ │ Regexes \u0026amp; Converters • Test Data Synthesizers │ └───────────────────────────┬────────────────────────────┘ │ Raw Implementation ▼ ┌────────────────────────────────────────────────────────┐ │ HUMAN ENGINEER │ │ Adversarial Audit • Failure Mode Verification │ │ Physical Profile (CPU/Mem/IO) • Production Deployment │ └────────────────────────────────────────────────────────┘ Treat the AI as an ultra-fast, junior typographical assistant. Let it generate the boilerplate structs, write the initial regex pattern, or scaffold the unit test stubs.\nKeep ownership of the boundaries:\nNever accept generated code you cannot mentally execute in your head. (See our architectural breakdown in The Four-Layer AI Coding Stack). Never deploy concurrent code without defining its physical saturation limits. Never let an AI assistant grade its own work. Maintain strict, independent test assertions, as detailed in Why I Don\u0026rsquo;t Let AI Grade Its Own Work. Recommended Learning Paths \u0026amp; Systems Foundations If you want to explore the deeper mechanics beneath the Renaissance Developer stack, check out these earlier deep dives:\nSilicon \u0026amp; Hardware Contracts: Read The ISA Is a Contract, Not a Benchmark to see how memory ordering and decoder silicon shape execution limits. Distributed State \u0026amp; Resilience: Read Sagas for Microservice Transactions and GitOps Reality with ArgoCD to master failure recovery without distributed locks. Pragmatic AI Workflows: Read The Four-Layer AI Coding Stack and Agentic Coding Assistants for production-proven guardrails. What Comes Next Understanding the systems moat is only the first step. To operate effectively as a Renaissance Developer, you must change how your brain absorbs and filters technical information.\nWhen every API answer and code snippet is available in seconds, memorizing syntax creates an illusion of competence that quickly atrophies real engineering judgment. In Part 2 of this series, we explore the cognitive mechanics of software engineering:\nRead Part 2: Meta-Cognition for Software Engineers: How to Think in the Age of Instant Answers to understand how to separate perishable syntax from timeless primitives, and how to build internal mental runtime engines that never decay. ","permalink":"https://hanhpham.vercel.app/posts/the-renaissance-developer/","summary":"When generative AI commoditizes syntax into zero-cost token emissions, the developer\u0026rsquo;s moat shifts from typing code to systems synthesis: understanding how a single concurrent change ripples across database row locks, cache lines, and cloud bills.","title":"The Renaissance Developer: Why AI Makes Systems Thinking the Ultimate Moat"},{"content":"A team notices that their continuous integration (CI) pipeline takes four minutes to build and test a Go microservice container.\nAn engineer reads a blog post about Docker BuildKit\u0026rsquo;s remote cache backends. They add two flags to the docker buildx build command in GitHub Actions:\n- name: Build and push uses: docker/build-push-action@v5 with: cache-from: type=gha cache-to: type=gha,mode=max They commit the change, expecting build times to drop to thirty seconds.\nInstead, the pull request pipeline run duration increases from 4 minutes to 6 minutes.\nThe runner spends 85 seconds downloading a 3.2GB cache archive from GitHub\u0026rsquo;s internal cache storage, 40 seconds extracting compressed layer tarballs to disk, 15 seconds executing the build, and another 110 seconds compressing and uploading the updated cache over the network before the job finishes.\nThe team spent compute, bandwidth, and storage quotas to save fifteen seconds of compilation time.\nCaching is not an automatic speed improvement. In cloud CI environments, caching is an economic trade-off between network transit latency and local CPU compute power.\nThe 30-Second Architecture: Remote CI layer caching is only faster than a scratch build when the time to download and decompress the cache over network interfaces is strictly less than the time required for local CPU cores to compute those layers from source. On modern 8-core or 16-core cloud runners with local NVMe disks, compiling a Go, Rust, or C++ binary from scratch is often twice as fast as streaming multi-gigabyte layer caches over public cloud network pipes. Staff engineers optimize build economics: cache lightweight immutable dependencies (go.mod, node_modules), avoid pushing multi-gigabyte compiler output over network caches, and structure multi-stage Dockerfiles with cache mounts (RUN --mount=type=cache) on persistent self-hosted runners.\nThe Remote Cache Trap: Network I/O vs. CPU Compute To evaluate whether remote caching makes sense for your pipeline, measure the build duration mathematically:\n$$\\text{Time}_{\\text{cached}} = \\text{Download Latency} + \\text{Decompression Latency} + \\text{Build Evaluation} + \\text{Upload Latency}$$$$\\text{Time}_{\\text{scratch}} = \\text{Source Checkout} + \\text{Compiler Execution} + \\text{Image Assembly}$$A remote cache provides a net positive return if and only if:\n$$\\text{Time}_{\\text{cached}} \u003c \\text{Time}_{\\text{scratch}}$$BuildKit with type=gha: |== Download Cache (85s) ==|== Extract (40s) ==|== Compile (15s) ==|== Compress/Upload (110s) ==| Total Job Duration: 250 seconds Scratch Build on Modern Runner: |== Checkout (5s) ==|== go mod download (8s) ==|== Compile from Source (35s) ==| Total Job Duration: 48 seconds The Three Network Bottlenecks Runner Network Throughput: GitHub-hosted standard runners and ephemeral CI virtual machines typically experience bandwidth caps between 40 MB/s and 80 MB/s. Decompression Overhead: Decompressing large gzip or zstd tarballs containing hundreds of thousands of small build artifacts (e.g. node_modules or .cache/go-build) saturates the runner\u0026rsquo;s single vCPU core, creating a CPU bottleneck before compilation even begins. Cache Upload Tax: If you specify mode=max, BuildKit exports cache manifests for every intermediate stage, including build tools, compilers, and temporary test binaries. Compressing and uploading that payload across every branch burns runner minutes. Layer Ordering \u0026amp; Invalidation Mechanics When remote caching does provide value (e.g. heavy base operating system layers or complex Cgo dependencies), the most common failure mode is premature layer invalidation.\nDocker evaluates cache validity sequentially from top to bottom. Once a single layer invalidates, Docker discards the cache for every subsequent instruction in that stage.\nThe Naive Dockerfile (Cache-Busting Anti-Pattern) # Antipattern: Breaks layer caching on every code edit FROM golang:1.24-alpine WORKDIR /app # Copying everything first invalidates the cache on every commit! COPY . . # This line runs from scratch every single time, re-downloading all dependencies: RUN go mod download RUN go build -o server ./cmd/server In the Dockerfile above, editing a comment in README.md changes the checksum of COPY . .. Docker throws away the cached RUN go mod download layer and re-downloads every dependency from the public internet on every commit.\nThe Optimized Multi-Stage Dockerfile Structure instructions from lowest rate of change to highest rate of change, and use BuildKit cache mounts:\n# Syntax directive enables modern BuildKit features # syntax=docker/dockerfile:1.7 # Stage 1: Build environment FROM golang:1.24-alpine AS builder WORKDIR /src # 1. Install system build tools (Quarterly change) RUN apk add --no-cache git make ca-certificates # 2. Copy dependency manifests only (Weekly change) COPY go.mod go.sum ./ # 3. Download dependencies with persistent cache mount # RUN --mount=type=cache persists across builds on the same runner node! RUN --mount=type=cache,target=/go/pkg/mod \\ go mod download # 4. Copy application source code (Hourly change) COPY . . # 5. Compile binary with compiler cache mount # CGO_ENABLED=0 produces a statically linked binary RUN --mount=type=cache,target=/go/pkg/mod \\ --mount=type=cache,target=/root/.cache/go-build \\ CGO_ENABLED=0 GOOS=linux GOARCH=amd64 \\ go build -trimpath -ldflags=\u0026#34;-s -w\u0026#34; -o /bin/server ./cmd/server # Stage 2: Minimal Distroless Runtime FROM gcr.io/distroless/static-debian12:nonroot WORKDIR /app # Copy only the compiled binary and certificates from the builder stage COPY --from=builder /etc/ssl/certs/ca-certificates.crt /etc/ssl/certs/ COPY --from=builder /bin/server /app/server # Run as non-root user (ID 65532 is built into distroless) USER nonroot:nonroot EXPOSE 8080 ENTRYPOINT [\u0026#34;/app/server\u0026#34;] Why This Wins Dependency Isolation: go.mod is copied separately. Editing application code in ./cmd/server leaves lines 1 through 10 fully cached. Cache Mounts (--mount=type=cache): Unlike layer caching (which packages files into image layers and exports them over the network), BuildKit cache mounts keep /go/pkg/mod and /root/.cache/go-build directly on the runner\u0026rsquo;s local filesystem. Even when a code edit forces recompilation, the Go compiler uses local cached object files (.a) to re-link in seconds. Distroless Attack Surface Reduction: The final image contains zero package managers (apk, apt), zero shells (/bin/sh, /bin/bash), and zero build tools. The final image size drops from 400MB to 18MB, speeding up deployment registry pulls across your Kubernetes cluster. Technology Trade-Off Matrix Strategy Network Overhead CPU Overhead Setup Complexity Best Production Use Case No Cache (Scratch Build) Low (Only pulls base image and modules) High (Recompiles every package) Zero Small Go/Rust binaries on modern, multi-core cloud runners. GHA Cache (type=gha) High (Uploads and downloads multi-GB blobs) Low Low (GitHub action config) Lightweight frontend builds, Python wheels (\u0026lt;500MB). Registry Cache (type=registry) Very High (Pushes cache layers to ECR/GHCR) Low Medium (Registry IAM setup) Base OS images, machine learning training images, multi-architecture builds. Persistent Local NVMe Runners Zero (Local disk read/write) Minimal (Incremental compilation) High (Requires managing self-hosted runners) Enterprise mono-repos, high-frequency continuous delivery teams. The Staff-Level Decision Framework Benchmark Scratch vs. Cached Before Enabling Remote Flags: Run five builds from clean source without remote caching. If a scratch build takes under 60 seconds, do not enable remote layer caching. The network transfer and decompression penalty will make your pipeline slower, not faster. Cache Dependencies, Not Build Artifacts: Use remote caching only for packages downloaded from public registries (npm, pip, go modules). Avoid exporting intermediate compilation caches (.cache/go-build, target/) over the network unless running on a persistent local runner. Use Self-Hosted Ephemeral Runners with Shared Local Cache Volumes: If build times exceed 10 minutes on large mono-repos, stop optimizing network layer caches. Deploy Kubernetes-based runner controllers (e.g. Actions Runner Controller - ARC) with persistent NVMe SSD cache volumes mounted into the builder pod. You eliminate network transit entirely while enjoying sub-second incremental builds. ","permalink":"https://hanhpham.vercel.app/posts/docker-buildkit-caching-economics-in-ci/","summary":"Enabling remote layer caching in CI often slows builds down instead of speeding them up. Here is how network transit limits, cache thrashing, and local NVMe runners dictate the real economics of container builds.","title":"Docker BuildKit in CI: The Economics of Network I/O vs CPU Compilation"},{"content":"Open Compiler Explorer, paste four lines of standard C++ concurrency primitives, and compile them for both x86-64 and AArch64 with optimization flag -O3:\n#include \u0026lt;atomic\u0026gt; std::atomic\u0026lt;int\u0026gt; g_x{0}; void sync_operations() { g_x.store(1, std::memory_order_release); int val = g_x.load(std::memory_order_acquire); g_x.fetch_add(1, std::memory_order_relaxed); std::atomic_thread_fence(std::memory_order_seq_cst); } The resulting assembly reveals the divide immediately:\n# --- x86-64 (GCC 14.2) --- sync_operations(): mov DWORD PTR g_x[rip], 1 # release store: standard mov! mov eax, DWORD PTR g_x[rip] # acquire load: standard mov! lock xadd DWORD PTR g_x[rip], eax # atomic RMW requires LOCK prefix lock add QWORD PTR [rsp], 0 # seq_cst fence: locked stack no-op! ret # --- AArch64 (GCC 14.2) --- sync_operations(): adrp x0, g_x mov w1, 1 stlr w1, [x0, :lo12:g_x] # release store: specialized STLR ldar w2, [x0, :lo12:g_x] # acquire load: specialized LDAR mov w1, 1 ldadd w1, w2, [x0, :lo12:g_x] # LSE atomic add: in-silicon RMW dmb ish # seq_cst fence: inner-shareable barrier ret On x86, release and acquire operations look free: they compile down to standard mov instructions. On ARM64, the programmer must explicitly emit specialized stlr and ldar opcodes. But when you ask for a sequential consistency fence, the situation flips: ARM emits an explicit data memory barrier (dmb ish), while x86 serializes the memory controller via an expensive locked store-buffer drain (lock add QWORD PTR [rsp], 0 or mfence).\nThe two processors arrive at the same destination, but they make opposite promises to the programmer.\nThe Core Takeaway: The classic textbook war between RISC and CISC ended decades ago when both sides converged on out-of-order execution engines running internal micro-operations. In 2026, the instruction set architecture (ISA) no longer dictates execution throughput. Instead, the ISA survives in three specific places where it still bills you:\nAs an immutable software contract: Memory consistency defaults (x86 TSO vs ARM weak ordering), register counts, and ABI calling conventions that compilers cannot paper over. As a hardware budget trade-off: Silicon die area spent on decoded micro-op caches (x86) versus massive L1 instruction caches (Apple Silicon). As an instrument of legal sovereignty: Proprietary licensing litigation (Arm vs Qualcomm) versus royalty-free specification profiles (RISC-V RVA23). 1. The Categories Died First: Silicon Convergence The 1980s curriculum taught a binary classification: CISC processors (x86) used complex, variable-length instructions that performed memory-to-register arithmetic to maximize code density on expensive magnetic-core RAM; RISC processors (MIPS, SPARC, early ARM) used simple, fixed 4-byte instructions strictly separating computation from memory access to enable single-cycle pipelining.\nIn modern silicon, that textbook taxonomy is completely obsolete:\n1980s Textbook Pillar Modern Silicon Reality (2026) Exemplar Implementation \u0026ldquo;CISC executes complex instructions directly\u0026rdquo; x86 decodes variable-length macro-instructions into fixed RISC-like micro-operations (micro-ops) before execution. Intel Lion Cove, AMD Zen 5 decode into an internal superscalar RISC core. \u0026ldquo;RISC executes one simple opcode per cycle\u0026rdquo; Modern ARM cores fuse multiple instructions together and execute complex multi-cycle vector matrix operations. Apple Firestorm fuses cmp + b.cc, AES pairs, and processes 128-bit LDP/STP dual-memory ops. \u0026ldquo;CISC requires complex microcode ROMs\u0026rdquo; High-performance x86 cores bypass decoders entirely using massive decoded micro-op caches. AMD Zen 5 embeds a 6,000-entry micro-op cache feeding up to 12 instructions per cycle. \u0026ldquo;RISC instruction sets are minimal and simple\u0026rdquo; ARM64 manuals span thousands of pages covering vector extensions, pointer auth, and memory tagging. ARMv9 introduces SVE2, SME (Streaming SVE), PAC, and MTE. Modern high-performance cores look virtually identical from the execution units backward. They are out-of-order, superscalar engines with hundreds of physical registers, speculative branch predictors, and multi-ported load/store queues.\nThe differences that survive are concentrated in the frontend and the software contract.\nTwo divergent silicon investments: x86 pays die area for a 6K-entry micro-op cache to bypass length decoding; Apple Silicon pays die area for a 192KB L1 instruction cache to feed an 8-wide direct decoder.\n2. The Decoder Tax: Real, Workload-Shaped, and Already Paid For A common claim in tech commentary asserts that x86 is permanently hobbled by a \u0026ldquo;decoder tax\u0026rdquo;: because x86 instructions range from 1 to 15 bytes in length, the CPU cannot know where instruction N+1 begins until it parses the prefixes and opcode of instruction N.\nThis creates a serialization bottleneck. To build a wide decoder (such as AMD Zen 5\u0026rsquo;s dual 4-wide decode clusters or Intel Lion Cove\u0026rsquo;s 8-wide decoder), x86 must run speculative pre-decode arrays to guess instruction boundaries across 16-byte fetch windows.\nIn contrast, ARM64 instructions are strictly fixed at 4 bytes (32 bits), naturally aligned. An 8-wide decoder (like Apple Silicon) simply slices a 32-byte cache window into eight equal 4-byte chunks with zero length speculation.\nMeasuring the Tax: The Zen 5 Op Cache Experiment How much does length decoding actually cost in performance?\nThe clearest empirical measurement was published by Chips and Cheese, who used AMD model-specific register MSR 0xC0011021 on a Ryzen 9 9900X to physically disable the 6,000-entry micro-op cache, forcing the CPU to fetch and decode every x86 instruction from L1I through the x86 decoders:\nWorkload Op Cache Coverage (Normal) Score Delta (Op Cache Disabled) What the Delta Proves SPECint (Single-Thread) 85% to 92% -20.3% Large instruction footprints with high single-thread IPC hit the 4-wide decode bottleneck hard. SPECfp (Single-Thread) 78% to 88% -16.8% Floating point loops encounter frontend starvation when falling back to legacy decode. SPECint (SMT Active) 82% to 90% -4.9% Dual SMT threads interleave decode clusters, reducing decoder idle bubbles. SPECfp (SMT Active) 75% to 85% -0.82% Compute units become execution-bound; frontend decode speed ceases to matter. Cinebench 2024 (1T) 84.4% -13.5% Rendering engine loops exceed small L1I windows when op cache is bypassed. Cyberpunk 2077 83.5% -0.17% Memory latency and GPU scheduling dominate; decode speed has virtually zero impact. GTA V 77.0% 0.0% (No change) Frame delivery is completely bound by draw calls and DRAM latency. The Silicon Budget Trade-Off This benchmark data dismantles the myth that x86 is doomed by its decoders:\nThe tax is already paid in silicon: Modern x86 processors hit 80% to 90% op cache coverage on typical code. The CPU spends die area on a 6,000-entry micro-op cache so that it rarely executes the legacy variable-length decoders during critical loops. The trade is symmetric: x86 spends transistors on decoded micro-op caches and predecode arrays. Apple Silicon spends transistors on a massive 192KB L1 instruction cache (3x larger than x86) to offset ARM64\u0026rsquo;s lower code density, and accepts one taken branch per cycle where x86 handles two. Both approaches are rational engineering trades optimized for different markets: Apple optimized for client battery efficiency and wide low-clock throughput; Intel and AMD optimized for high-clock server density and 40 years of binary compatibility.\n3. The Contract the Compiler Cannot Paper Over While out-of-order execution engines absorb instruction encoding differences, the ISA defines a rigid contract with the compiler and the operating system.\nThe ISA in 2026: instruction encoding is largely absorbed by execution units; the contract and economics remain immutable.\n1. Memory Consistency Models: x86-TSO vs ARM Weak Ordering The most consequential difference between architectures is the memory consistency model:\nx86 Total Store Order (TSO): Hardware guarantees that memory stores become visible in strict program order. A store cannot pass an earlier store, and a load cannot pass an earlier load. Relaxed and acquire/release memory orders compile to regular mov instructions because the hardware enforces TSO by default. ARM Weak Ordering: The memory bus permits aggressive reordering of loads and stores. To establish an acquire/release relationship, compilers must explicitly emit specialized instructions (LDAR, STLR) or insert full memory barriers (DMB ISH). x86-64 Memory Contract (Strict TSO): Store A =====\u0026gt; Store B (Hardware guarantees visibility order: B never seen before A) ARM64 Memory Contract (Weakly Ordered): Store A ---\\ /--- Store B (Reordered by interconnect unless guarded by STLR or DMB) X Store B ---/ \\--- Store A The Rosetta 2 Hardware Trick When Apple built the M1 chip to run x86 binaries via Rosetta 2 translation, they encountered a severe architectural challenge: emulating x86 TSO semantics on a weakly-ordered ARM64 core requires injecting memory barrier instructions (DMB) around almost every translated memory access. Doing so in software degrades translation performance by 30% to 40%.\nApple\u0026rsquo;s solution was pure engineering audacity: they added an undocumented hardware register bit to their CPU cores.\nWhen the macOS kernel schedules an x86 binary running under Rosetta 2, it flips an internal CPU control register: ACTLR_EL1.TSO = 1\nThis hardware switch puts the Apple Silicon core into x86 Total Store Order mode. The memory controller\u0026rsquo;s store queue begins enforcing x86 store ordering rules directly in silicon.\nThe cost? Benchmarks measuring Apple Silicon with hardware TSO enabled show a ~9% performance penalty on multi-threaded SPEC workloads due to conservative store-buffer draining. But that 9% silicon tax was vastly cheaper than the 40% penalty of software barrier emulation.\nThis is the ultimate proof that the ISA matters: an entire physical hardware mode exists in Apple Silicon solely to honor another architecture\u0026rsquo;s memory contract.\nThe Sunset of Rosetta 2 Apple published a developer update confirming that general-purpose Rosetta 2 support will sunset after macOS 27, leaving only a narrow compatibility layer for legacy games. The bridge was built to migrate the ecosystem; once the software contract migrated natively to AArch64, the hardware justification for maintaining x86 TSO modes disappeared.\n2. Register Pressure and Code Density General Purpose Registers: Classic x86-64 exposes only 16 architectural GPRs. ARM64 and RISC-V expose 31 and 32 GPRs respectively. In tight mathematical loops, 16 registers force the compiler to spill variables to the stack cache, increasing load/store instruction traffic. Intel\u0026rsquo;s APX (Advanced Performance Extensions) finally expands x86 to 32 GPRs with 3-operand instruction syntax, acknowledging the register-pressure penalty. Code Size \u0026amp; Instruction Cache Footprint: Because x86 instructions are variable-length (1 to 15 bytes), CISC code is dense. A single x86 instruction can encode complex scaled addressing and arithmetic: ADD [rax + rbx*4 + 0x20], edx On ARM64, the same operation requires separate address calculation, load, add, and store instructions. Consequently, ARM64 .text sections are commonly 10% to 25% larger than equivalent x86-64 binaries, requiring larger L1 instruction caches to maintain parity. 4. The ISA as a Legal and Economic Instrument Beyond silicon and compilers, the modern ISA operates as a legal instrument dictating who is allowed to manufacture microprocessors.\n1. The Three Licensing Models The x86 Closed Duopoly: Intel and AMD maintain a perpetual cross-licensing patent treaty. No third party can license x86 to build custom server or mobile silicon. If a cloud hyperscaler wants custom silicon, x86 is legally unavailable. The ARM IP Licensing Model: Arm Ltd licenses both pre-designed CPU cores (Cortex-X, Neoverse) and Architectural Licenses (allowing companies like Apple and Qualcomm to design custom microarchitectures from scratch). However, the ongoing legal dispute between Arm and Qualcomm over the Nuvia/Oryon core acquisition demonstrates the commercial friction of proprietary ISA governance. RISC-V Open Specification: RISC-V is an open standard governed by a Swiss entity (RISC-V International). The specification is free and cannot be revoked by export controls. Anyone can design a RISC-V core without paying royalties or requesting permission. 2. RISC-V\u0026rsquo;s Real World in 2026: The Accelerator Sub-Processor Tech evangelists often predict that RISC-V will displace x86 and ARM in laptops and servers overnight. That view misunderstands software distribution economics.\nThe real challenge for application-class RISC-V is software fragmentation. Because RISC-V was designed as a modular base with dozens of optional extensions, early Linux distributions faced an explosion of incompatible binary targets. The RVA23 Profile Standard was created specifically to freeze a unified instruction baseline for commercial operating systems.\nWhere RISC-V is winning today is not in primary host CPUs, but inside accelerator sub-processors:\nModern enterprise GPUs, AI accelerators, and network interface cards (SmartNICs) embed dozens of microcontroller cores to manage firmware, power domains, thermal loops, and memory telemetry. NVIDIA completely replaced its proprietary Falcon control microcontrollers with custom NV-RISCV32 and NV-RISCV64 cores across its GPU line. Western Digital and Seagate ship billions of storage controllers driven by RISC-V. In these environments, binary compatibility with Windows or Debian is irrelevant; zero license fees and complete architectural freedom dominate.\nMeasuring Your Own Frontend Boundness If you want to know whether your production workload actually cares about instruction decoding or frontend stalls, stop reading architectural debates and measure it directly with Linux perf:\n# Measure topdown frontend vs backend execution metrics perf stat --topdown -a -e cycles,instructions,frontend_retired.any_ds -- sleep 10 On Linux systems with TopDown profiling support (Intel Ice Lake / Sapphire Rapids and AMD Zen 4/5), perf splits execution cycles into four primary buckets:\n[Pipeline Slots] ├── Frontend Bound (Stalled on instruction fetch, decode, or uop cache miss) ├── Bad Speculation (Discarded work from branch mispredictions) ├── Backend Bound (Stalled on memory latency, L3 cache, or execution ports) └── Retiring (Useful work completed) In 90% of real-world distributed backends (Go services, database query engines, Java runtimes), the profiler reveals that Backend Bound memory stalls account for 50% to 70% of cycles, while Frontend Bound decode stalls sit under 10%.\nUnless your service is spending its life thrashing a 2MB instruction cache inside an instruction-heavy binary, the instruction set architecture is not your bottleneck. Memory bandwidth, cache hierarchy, and concurrency design dictate your latency.\nSources \u0026amp; Technical References Blem, Menon, Sankaralingam (2013): \u0026ldquo;Power Struggles: Revisiting the RISC vs. CISC Debate on Contemporary ARM and x86 Architectures\u0026rdquo; (HPCA 2013). Empirical proof that ISA differences have negligible impact on performance compared to microarchitectural implementation. Chips and Cheese (2024): \u0026ldquo;Disabling Zen 5\u0026rsquo;s Op Cache and Exploring its Clustered Decoder\u0026rdquo;. Primary benchmark data measuring Zen 5 performance drop with op cache disabled via MSR 0xC0011021. Dougall Johnson (2021–2024): \u0026ldquo;Apple Silicon M1/Firestorm Microarchitecture Documentation\u0026rdquo;. Reverse-engineered measurements of 8-wide decode, coalesced retire queues, and instruction fusion. AMD Corporation (2024): \u0026ldquo;Software Optimization Guide for AMD Family 1Ah Processors (Zen 5)\u0026rdquo;. Architecture specs for 6K-entry op cache and dual decode clusters. Intel Corporation (2024): \u0026ldquo;Intel 64 and IA-32 Architectures Optimization Reference Manual\u0026rdquo;. Lion Cove 8-wide decoder and Decoded Stream Buffer (DSB) specs. Apple Developer Documentation (2025): \u0026ldquo;Rosetta 2 Transition \u0026amp; Support Lifecycle Note\u0026rdquo;. Developer schedule confirming Rosetta 2 sunset. ","permalink":"https://hanhpham.vercel.app/posts/isa-as-a-contract-risc-vs-cisc/","summary":"Four decades of micro-op convergence absorbed the classic textbook differences between RISC and CISC, but the instruction set still bills you in three places you cannot optimize away: memory ordering contracts, frontend silicon budgets, and ecosystem sovereignty.","title":"The ISA Is a Contract, Not a Benchmark: What the Instruction Set Still Costs You in 2026"},{"content":"Every software engineer eventually hits a wall with static analysis. You write a linter rule, build an automated vulnerability scanner, or write a compiler pass, and you wonder: why can\u0026rsquo;t a tool definitively tell me whether this loop terminates, whether this pointer ever dereferences null, or whether two arbitrary functions produce the same output? The answer was proved in 1936, before physical computers even existed.\nThe Core Takeaway: In his landmark 1936 paper \u0026ldquo;On Computable Numbers, with an Application to the Entscheidungsproblem\u0026rdquo;, 24-year-old Alan Turing answered David Hilbert\u0026rsquo;s decision problem with a definitive NO. To do so, he invented the mathematical blueprint for modern computers (the Universal Turing Machine), proved that the Halting Problem is undecidable via a fatal self-referential diagonal contradiction, and demonstrated that mathematical truth outstrips mechanical computation.\n1. The Context: Hilbert\u0026rsquo;s Dream of Total Automation In 1928, mathematician David Hilbert challenged the mathematical world to resolve three foundational questions about formal axiomatic systems:\nIs mathematics complete? Can every true statement be proved from axioms? (Answered NO by Kurt Gödel in 1931: any consistent system capable of arithmetic contains true statements that cannot be proven within the system). Is mathematics consistent? Can we prove that the axioms never lead to a contradiction like \\( 0 = 1 \\)? (Answered NO by Gödel\u0026rsquo;s Second Theorem: a system cannot prove its own consistency). Is mathematics decidable? (Das Entscheidungsproblem, or The Decision Problem): Is there an effective mechanical procedure (an algorithm) that takes any statement in first-order logic and decides, in a finite number of steps, whether it is universally valid?\nHilbert believed mathematics was fully mechanized. If the answer to the third question was yes, all mathematical inquiry could be handed over to a machine: feed in a conjecture (like the Riemann Hypothesis or Goldbach\u0026rsquo;s Conjecture), crank the algorithmic handle, and receive a definitive True or False.\nTo disprove Hilbert, Turing had to first answer a question nobody had formalized: what is an algorithm?\n2. Formalizing the \u0026ldquo;Computer\u0026rdquo;: The a-Machine In 1936, the word \u0026ldquo;computer\u0026rdquo; referred to a human clerk performing calculations on paper. Turing deconstructed what that human actually does down to physical primitives:\nA clerk works on paper: Turing abstracted this into an infinite 1-dimensional tape divided into discrete squares. A clerk can only inspect a limited number of symbols at one glance: Turing restricted the machine head to scanning one square at a time. A clerk has a finite number of distinct \u0026ldquo;states of mind\u0026rdquo;: Turing defined a finite set of internal states \\( Q = \\{q_0, q_1, \\dots, q_k\\} \\). A clerk moves between cells, alters symbols, and updates their train of thought based on rules: Turing formalized this as a finite transition function: $$\\delta: (q_i, s_j) \\longrightarrow (q_{new}, s_{write}, \\text{Direction})$$ ... | Blank | 1 | 0 | 1 | 1 | 0 | Blank | ... \u0026lt;-- Infinite Memory Tape ^ | [ Read/Write Head ] +---------+ | State q | Transition Rule: (State q, Read 0) -\u0026gt; (Write 1, Move R, State q\u0026#39;) +---------+ Turing termed this an a-machine (automatic machine). Despite having only four basic physical actions (read, write, move left/right, change state), this abstract construct can simulate any algorithm ever conceived.\nTuring then classified real numbers:\nA real number in the interval \\([0, 1]\\) is computable if its digits can be printed sequentially on the tape by an a-machine. Machines that run indefinitely printing an infinite sequence of digits are called circle-free. Machines that halt, crash, or enter infinite loops without printing digits are called circular. 3. Code Is Data: Description Numbers \u0026amp; The Universal Machine Because a machine\u0026rsquo;s transition rules are finite, its complete specification can be written as a table of text.\nStandard Descriptions (S.D.): Turing assigned a fixed character encoding to every state and transition. Description Numbers (D.N.): By mapping characters to integers, every unique Turing machine compresses into a single, finite positive integer. Machine M ---\u0026gt; Standard Description (S.D.) ---\u0026gt; Unique Integer: D.N.(M) This led to two groundbreaking breakthroughs:\nConsequence 1: Programs are Countable Integers Because every Turing machine corresponds to a unique integer \\( D.N. \\in \\mathbb{N} \\), the set of all possible computer programs is countably infinite: $$M_1, M_2, M_3, M_4, \\dots$$There are no more computer programs in existence than there are whole numbers.\nConsequence 2: The Universal Turing Machine (\\(\\mathcal{U}\\)) If a program is just an integer, a machine can read another program as data.\nTuring designed a single, specific machine, \\(\\mathcal{U}\\), which accepts two inputs on its tape: the Description Number \\( D.N.(M) \\) of any arbitrary machine \\( M \\), and an input string \\( x \\). Machine \\(\\mathcal{U}\\) reads the rules of \\( M \\) from the tape and simulates its execution step by step: $$\\mathcal{U}(D.N.(M), x) \\equiv M(x)$$This is the invention of the stored-program computer. Before Turing, hardware was built to do one task (a cash register added; a loom wove). Turing showed that a single physical piece of hardware could execute any arbitrary software program.\n4. The Proof: Cantor\u0026rsquo;s Diagonalization Applied to Machines With the Universal Machine established, Turing addressed the core question: can a machine determine whether another machine will run forever or stall?\nToday this is known as the Halting Problem (Turing formulated it as determining whether a machine is \u0026ldquo;circle-free\u0026rdquo;):\nThe Hypothesis (Proof by Contradiction) Assume there exists an algorithm: a Turing machine \\(\\mathcal{D}\\) that can inspect any machine Description Number \\( n \\) and decide whether it is circle-free: $$\\mathcal{D}(n) = \\begin{cases} \\text{True} \u0026 \\text{if machine } M_n \\text{ will print digits forever (circle-free)} \\\\ \\text{False} \u0026 \\text{if machine } M_n \\text{ halts, stalls, or loops without output (circular)} \\end{cases}$$Constructing the Machine \\(\\mathcal{H}\\) If \\(\\mathcal{D}\\) exists, we can construct a new machine \\(\\mathcal{H}\\) that generates a complete list of all computable real numbers:\nIncrement an integer \\( n = 1, 2, 3, \\dots \\) Run \\(\\mathcal{D}(n)\\). If \\(\\mathcal{D}(n) = \\text{False}\\), discard \\( n \\). If \\(\\mathcal{D}(n) = \\text{True}\\), invoke the Universal Machine \\(\\mathcal{U}\\) to compute the digits of \\( M_n \\). This creates an ordered, exhaustive matrix of all computable real numbers:\n$$\\begin{aligned} M_1: \u0026\\quad 0.\\mathbf{d_{1,1}} \\; d_{1,2} \\; d_{1,3} \\; d_{1,4} \\dots \\\\ M_2: \u0026\\quad 0.d_{2,1} \\; \\mathbf{d_{2,2}} \\; d_{2,3} \\; d_{2,4} \\dots \\\\ M_3: \u0026\\quad 0.d_{3,1} \\; d_{3,2} \\; \\mathbf{d_{3,3}} \\; d_{3,4} \\dots \\\\ M_4: \u0026\\quad 0.d_{4,1} \\; d_{4,2} \\; d_{4,3} \\; \\mathbf{d_{4,4}} \\dots \\\\ \\vdots \\end{aligned}$$The Diagonal Construction (\\(\\beta\\)) Borrowing Georg Cantor\u0026rsquo;s 1891 diagonal argument, Turing defines a new number \\(\\beta = 0.\\beta_1 \\beta_2 \\beta_3 \\dots\\) by systematically inverting the diagonal elements: $$\\beta_k = 1 - d_{k,k}$$If the \\( k \\)-th digit of machine \\( M_k \\) is \\( 0 \\), make \\( \\beta_k = 1 \\). If it is \\( 1 \\), make \\( \\beta_k = 0 \\).\nM1: [0] 1 0 1 ... -\u0026gt; flip to 1 M2: 1 [1] 0 0 ... -\u0026gt; flip to 0 M3: 0 0 [0] 1 ... -\u0026gt; flip to 1 ... Beta: 1 0 1 ... The Fatal Contradiction Now evaluate two questions about \\(\\beta\\):\nIs \\(\\beta\\) computable?\nYes. We have an exact mechanical procedure to find every digit: to find digit \\( k \\), run \\(\\mathcal{H}\\) to find the \\( k \\)-th circle-free machine, compute its \\( k \\)-th digit \\( d_{k,k} \\), and flip it (\\( 1 - d_{k,k} \\)). Because \\(\\beta\\) is computable, it must be produced by some machine in our list: let us say machine \\( M_m \\). What is the \\( m \\)-th digit of \\(\\beta\\)? By definition of \\(\\beta\\): \\( \\beta_m = 1 - d_{m,m} \\). But because \\(\\beta\\) is the number computed by \\( M_m \\), its \\( m \\)-th digit must be: \\( \\beta_m = d_{m,m} \\). Equating the two: $$d_{m,m} = 1 - d_{m,m} \\implies 2 \\cdot d_{m,m} = 1$$No binary digit in \\(\\{0, 1\\}\\) can satisfy this equation.\nThe premise is false. The decision machine \\(\\mathcal{D}\\) cannot exist.\nThere is no general algorithm that can determine whether an arbitrary program will halt or run forever.\n5. The Application: Destroying the Entscheidungsproblem Having proved that the Halting Problem is uncomputable, Turing applied this result directly to Hilbert\u0026rsquo;s Entscheidungsproblem.\ngraph LR TM[\u0026#34;Arbitrary Turing Machine M\u0026#34;] --\u0026gt; Formula[\u0026#34;Construct First-Order Formula Un(M)\u0026#34;] Formula --\u0026gt; Solver[\u0026#34;Hypothetical Decision Algorithm E\u0026#34;] Solver --\u0026gt; Halting[\u0026#34;Decides whether M halts\u0026#34;] Halting -.-\u0026gt; Contradiction[\u0026#34;Contradiction: Halting is Undecidable!\u0026#34;] Encoding Execution in Logic: For any Turing machine \\( M \\), Turing showed how to mechanically construct a single, finite formula in first-order predicate logic, denoted \\( \\text{Un}(M) \\), that describes: The initial blank tape configuration. The legal transitions of \\( M \\). The assertion: \u0026ldquo;Machine \\( M \\) eventually prints the symbol 0.\u0026rdquo; The Logical Equivalence: Turing proved: $$\\text{Machine } M \\text{ eventually prints 0} \\iff \\text{Formula } \\text{Un}(M) \\text{ is provable in first-order logic}$$ The Conclusion: If Hilbert\u0026rsquo;s decision procedure existed, an algorithm could take \\( \\text{Un}(M) \\) and decide its validity in finite time. But doing so would solve the Halting Problem, which was just proven mathematically impossible. Therefore, first-order logic is undecidable. There is no algorithm that can determine the truth of all mathematical statements.\n6. What This Means for Modern Software Engineering Turing\u0026rsquo;s 1936 paper is not ancient history: it draws the outer boundary of what every compiler, linter, and security analyzer can ever achieve:\nRice\u0026rsquo;s Theorem (1953): Any non-trivial semantic property of a program (e.g. \u0026ldquo;does this function ever leak memory?\u0026rdquo;, \u0026ldquo;is this API route vulnerable to SQL injection?\u0026rdquo;, \u0026ldquo;does this routine return 403?\u0026rdquo;) is undecidable. Why Static Analysis Employs Approximations: Linters and scanners cannot be simultaneously sound (no false negatives) and complete (no false positives). Every security tool is mathematically forced to make trade-offs: either flag benign code (false alarms) or miss real bugs (false sense of security). Turing Completeness as an Attack Surface: Whenever a configuration format (YAML, CSS, PDF font engines, BPF, smart contracts) accidentally becomes Turing complete, verifying its safety in advance becomes impossible. Turing did not merely find a boundary in mathematics: by proving what algorithms cannot do, he built the first complete description of what all computers can do.\n","permalink":"https://hanhpham.vercel.app/posts/alan-turing-computable-numbers-entscheidungsproblem/","summary":"A step-by-step deconstruction of Alan Turing\u0026rsquo;s 1936 paper: how formalizing the mechanical computer and adapting Cantor\u0026rsquo;s diagonal argument proved that algorithmic omniscience is mathematically impossible.","title":"Can Every Problem Be Solved by an Algorithm? Alan Turing's 1936 Proof"},{"content":"Your team adopts GitOps with ArgoCD. The pitch from conferences and vendor blogs is seductive:\n\u0026ldquo;Everything is code. Git is your single source of truth. If a production release fails, just revert the Git commit or click \u0026lsquo;Rollback\u0026rsquo; in the ArgoCD UI, and the cluster self-heals back to safety.\u0026rdquo;\nYou deploy version 2.4.0 of your core billing service. The GitOps pipeline detects the commit, runs a pre-sync Kubernetes Job to execute database migrations, and updates the deployment manifest.\nTen minutes later, a critical application bug surfaces in version 2.4.0: an edge-case calculation causes payment timeouts. The on-call engineer navigates to the ArgoCD dashboard and clicks Rollback to v2.3.0.\nArgoCD terminates the v2.4.0 pods and redeploys the v2.3.0 container image.\nInstantly, all v2.3.0 pods enter CrashLoopBackOff:\n2026-09-30T09:14:22.819Z [FATAL] server.go:142: failed to initialize account repository: pq: column \u0026#34;billing_account_status\u0026#34; does not exist in table \u0026#34;accounts\u0026#34; goroutine 1 [running]: main.setupDatabase(0x104b2a0) /src/repository/accounts.go:88 +0x31a main.main() /src/cmd/server/main.go:44 +0x182 The database migration job in v2.4.0 renamed billing_account_status to status_code and dropped the old column. The v2.3.0 application code cannot find the column it expects. Automated rollback did not save the system; it transformed a partial application bug into complete catastrophic cluster downtime.\nGitOps is an exceptional tool for managing stateless declarations. When applied blindly to stateful database dependencies, it creates dangerous failure modes.\nThe 30-Second Architecture: Git commits can be reverted in milliseconds; database mutations cannot. Automated GitOps rollbacks fail because Kubernetes containers are ephemeral, while databases are persistent and append-only. When a pre-sync migration modifies schema destructively, reverting the Git deployment manifest deploys legacy application code against a mutated database. Staff engineers decouple application deployment from database migration lifecycle: schema migrations must strictly follow the Expand/Contract pattern, guaranteeing that database schema version N is backward-compatible with application versions N and N-1 simultaneously before any GitOps sync triggers.\nThe Declarative Fantasy vs. Stateful Reality The core premise of GitOps is that Cluster State = Git State.\n[ Git Repository: Desired State ] \u0026lt;--- (ArgoCD Reconciliation Loop) ---\u0026gt; [ Live Kubernetes Cluster ] This model works reliably for stateless constructs:\nDeployments, ReplicaSets, and Pods. Services, Ingress objects, and Gateway APIs. ConfigMaps and Secrets. If a ConfigMap typo breaks a service, reverting the Git commit updates the ConfigMap, restarts the pod, and resolves the issue cleanly.\nDatabases violate the GitOps premise because the database schema is not a Kubernetes manifest.\nWhen you bundle SQL migrations into an ArgoCD PreSync hook:\nGitOps reconciles the repository. The PreSync Job connects to PostgreSQL, MySQL, or CockroachDB and executes DDL statements (ALTER TABLE, DROP COLUMN, ADD CONSTRAINT). The database state transitions irrevocably to version 2. The deployment updates pods to version 2. If version 2 fails, Git cannot \u0026ldquo;revert\u0026rdquo; the database. Reverting a Git commit simply reverts the container image pointer back to version 1. The database remains at version 2. Unless you have engineered backward compatibility into your database schema, version 1 code crashes immediately on boot.\nThe Sync Wave \u0026amp; Hook Ordering Traps To coordinate complex multi-tier applications, ArgoCD provides Sync Waves and Resource Hooks. While powerful, they introduce three systems-level failure modes at scale:\n1. The Deadlock of the Failed PreSync Job ArgoCD sync waves execute in strictly sequential phases:\napiVersion: batch/v1 kind: Job metadata: name: schema-migration-v2-4-0 annotations: argocd.argoproj.io/hook: PreSync argocd.argoproj.io/hook-delete-policy: BeforeHookCreation If the migration Job fails (e.g. an ALTER TABLE query times out acquiring an exclusive table lock due to long-running analytical queries), ArgoCD halts the sync operation.\nThe failure mode: The deployment enters a frozen state.\nThe old pods remain running. The new pods are blocked from deploying. The failed Job locks the sync queue. If an engineer tries to push an emergency hotfix to Git, ArgoCD refuses to sync because the previous PreSync phase never succeeded. Manual intervention (argocd app terminate-op) is required to break the lock. 2. The CRD Race Condition When managing platform components (cert-manager, Prometheus Operator, Istio, External Secrets) via GitOps, teams frequently bundle CustomResourceDefinitions (CRDs) alongside the Custom Resources (CRs) that depend on them:\nmonitoring/ ├── prometheus-operator-crd.yaml └── production-servicemonitor.yaml During git synchronization, ArgoCD sends both objects to the Kubernetes API server simultaneously. The API server accepts the CRD, but the internal schema validator has not finished establishing the OpenAPI v3 validation schema before the ServiceMonitor manifest arrives.\nThe API server rejects the Custom Resource with:\nerror: unable to recognize \u0026#34;production-servicemonitor.yaml\u0026#34;: no matches for kind \u0026#34;ServiceMonitor\u0026#34; in version \u0026#34;monitoring.coreos.com/v1\u0026#34; The GitOps sync fails on clean, valid code purely due to asynchronous API registration latency.\nTechnology Trade-Off Matrix Strategy Operational Advantage Hidden Failure Mode Production Sweet-Spot In-Band PreSync Hooks (Default) Migrations and code deploy in a single Git commit; zero external tooling. Automated Git rollbacks crash legacy code; migration failures freeze the sync queue. Prototypes, staging environments, single-developer projects. Decoupled Pipeline Migrations Migrations execute in separate CI pipeline before GitOps sync triggers. Requires external pipeline orchestration outside of pure GitOps. Medium-scale services with mature CI/CD gating. Expand/Contract Schema Evolution Zero-downtime rollouts and instantaneous, safe rollbacks at any hour. Requires writing multiple backward-compatible PRs for a single database refactor. Mandatory for enterprise, high-availability, business-critical systems. The Staff-Level Decision Framework To build a reliable delivery architecture, you must decouple the database lifecycle from application deployment.\n+-----------------------------------------------------------------------------------+ | THE EXPAND / CONTRACT PROTOCOL | +-----------------------------------------------------------------------------------+ | Phase 1: Expand | Add new column as nullable; dual-write in application code.| | | Both v1 and v2 run safely against this schema. | |----------------------+------------------------------------------------------------| | Phase 2: Migrate | Backfill existing records asynchronously via worker jobs. | |----------------------+------------------------------------------------------------| | Phase 3: Contract | Switch application reads to new column; drop old column | | | only after v1 code is 100% retired from the cluster. | +-----------------------------------------------------------------------------------+ 1. The Three-Phase Expand/Contract Rule Never rename or drop a column in a single migration. Every destructive schema change must span three distinct deployments:\nStep 1 (Expand): Add the new column status_code alongside the old column billing_account_status. Make the new column nullable. Deploy application code that reads from billing_account_status but writes to both columns. Step 2 (Backfill): Run an asynchronous background script to copy historical data from the old column to the new column. Step 3 (Contract): Deploy application code that reads and writes exclusively from status_code. Only after this version is stable in production and previous versions are retired do you issue a final migration to drop the old column. If a bug occurs in Step 3, you can safely roll back to Step 2 or Step 1: the database supports both application versions simultaneously.\n2. Solving the CRD Race with Negative Sync Waves To prevent API server schema rejection during operator deployments, use negative sync waves to force CRDs to establish before Custom Resources compile:\n# In CRD definition: apiVersion: apiextensions.k8s.io/v1 kind: CustomResourceDefinition metadata: name: servicemonitors.monitoring.coreos.com annotations: argocd.argoproj.io/sync-wave: \u0026#34;-2\u0026#34; --- # In Custom Resource: apiVersion: monitoring.coreos.com/v1 kind: ServiceMonitor metadata: name: api-metrics annotations: argocd.argoproj.io/sync-wave: \u0026#34;0\u0026#34; ArgoCD guarantees that wave -2 is completely applied and registered with the API server before wave 0 begins evaluation.\n3. Configure ignoreDifferences for Dynamic Controllers When Kubernetes controllers mutate resources dynamically, ArgoCD detects unexpected drift and triggers continuous sync loops.\nExplicitly configure ignoreDifferences in your ArgoCD Application manifest:\nspec: ignoreDifferences: # Ignore replica scaling managed dynamically by HPA - group: apps kind: Deployment jsonPointers: - /spec/replicas # Ignore webhook-injected annotations and certificates - group: \u0026#34;\u0026#34; kind: Service jsonPointers: - /metadata/annotations/service.beta.kubernetes.io~1aws-load-balancer-arn ","permalink":"https://hanhpham.vercel.app/posts/gitops-argocd-database-migration-reality/","summary":"GitOps treats application deployments as stateless Git diffs, but production backends run on stateful databases. Here is why automated Git rollbacks break under schema migrations, and how to decouple delivery safely.","title":"GitOps Reality with ArgoCD: When Declarative State Fights the Database"},{"content":"Your Go microservice has an established latency SLO: p99 response times must remain below 30 milliseconds.\nDuring a normal traffic day, average CPU utilization sits comfortably at 22%. Yet your monitoring alerts trigger: p99 latency has spiked to 320 milliseconds. Endpoints that typically execute in 4 milliseconds are stalling.\nAn on-call engineer checks Datadog, notes that CPU utilization is only 22%, assumes the database is slow, and pages the DBA team. The database team verifies that query latency is sub-millisecond. In confusion, someone doubles the container CPU limit from 2000m to 4000m. The latency spikes decrease slightly, but return during the next traffic burst.\nMeanwhile, another service replica abruptly disappears from the cluster without printing a single stack trace or error log. Kubernetes status reports OOMKilled (Exit Code 137).\nBoth failures originate from the exact same architectural misunderstanding: treating Kubernetes CPU and memory limits as generic resource boundaries rather than Linux kernel cgroup primitives.\nThe 30-Second Architecture: Kubernetes CPU limits do not cap clock frequency; they enforce rigid Completely Fair Scheduler (CFS) quotas in 100ms period windows. Multi-threaded runtimes (Go, JVM, Node) handling concurrent bursts consume their entire period quota in a fraction of a wall-clock second, causing the kernel to freeze the process for the remainder of the period. For latency-critical microservices, hard CPU limits are an anti-pattern: set deterministic CPU requests, leave CPU limits unset (limits.cpu: null), and use Linux CPU shares (cpu.weight) to allocate contested compute. For memory, enforce hard container limits and pair them with Go\u0026rsquo;s GOMEMLIMIT at 85% capacity to trigger aggressive garbage collection before the kernel invokes oom_score_adj.\nLinux CFS Quotas Under the Hood To understand why a pod running at 20% average CPU experiences 300ms latency spikes, you must inspect how the Linux kernel throttles cgroups.\nIn Kubernetes, setting resources.limits.cpu: \u0026quot;1000m\u0026quot; translates directly into two Linux cgroup parameters:\ncpu.cfs_period_us: The quota measurement window, configured globally by Kubelet to 100,000 microseconds (100 milliseconds). cpu.cfs_quota_us: The total CPU runtime allocated to the container within that window. For 1000m (1 core), this is set to 100,000 microseconds. For 2000m (2 cores), this is 200,000 microseconds. The critical mechanics: quota is cumulative across all threads, but the period window is bound to wall-clock time.\nPeriod Window: 100ms (Wall Clock) Quota Allocated: 200ms CPU runtime (equivalent to limits.cpu: 2000m) Scenario: Service receives 10 concurrent requests. Go runtime schedules work across 10 active OS threads. Thread 1: [ 20ms work ] Thread 2: [ 20ms work ] Thread 3: [ 20ms work ] Thread 4: [ 20ms work ] Thread 5: [ 20ms work ] --\u0026gt; Combined CPU runtime: 200ms consumed in 20ms of wall time! Thread 6: [ 20ms work ] Thread 7: [ 20ms work ] Thread 8: [ 20ms work ] Thread 9: [ 20ms work ] Thread 10: [ 20ms work ] Wall Clock: |=== 20ms Active ===|================ 80ms Kernel Freeze ================| ^ ^ ^ 0ms 20ms (Quota Exhausted) 100ms (Next Period) In the diagram above:\nThe container has a limit of 2 cores (200ms quota per 100ms period). Ten goroutines process work concurrently. In just 20 milliseconds of wall-clock time, the 10 threads accumulate 200 milliseconds of compute time. The cgroup quota is completely spent. The Linux kernel suspends the entire cgroup for the remaining 80 milliseconds. To external clients, the application stops responding. In-flight TCP packets sit unacknowledged in the kernel socket receive buffer. When the next 100ms period opens, the kernel unsuspends the threads, only for the burst to exhaust the quota again.\nBecause the container was active for only 20ms out of every 100ms, standard metrics tools report: $$\\text{Average CPU Utilization} = \\frac{200\\text{ms runtime}}{1000\\text{ms sample interval}} = 20\\%$$The service is throttled 80% of the time, yet the dashboard shows green.\nHow to Detect It Query Prometheus for raw CFS throttling counters rather than average CPU usage:\n# Percentage of periods spent throttled sum(rate(container_cpu_cfs_throttled_periods_total{container=\u0026#34;api\u0026#34;}[5m])) / sum(rate(container_cpu_cfs_periods_total{container=\u0026#34;api\u0026#34;}[5m])) * 100 If this metric exceeds 5% on a synchronous service, your p99 latency spikes are generated by the Linux kernel, not application code.\nThe Go Runtime Multiplier: The GOMAXPROCS Trap If you deploy a Go microservice to Kubernetes without explicitly configuring GOMAXPROCS, the problem amplifies exponentially.\nBy default, the Go runtime calls runtime.NumCPU() on boot to determine how many operating system scheduler threads (\\(M\\)) to spin up. In containerized environments, runtime.NumCPU() reads /sys/devices/system/cpu/online from the host node.\nIf your pod runs on an AWS c6i.16xlarge worker node with 64 physical CPU cores, the Go runtime defaults to: $$\\text{GOMAXPROCS} = 64$$If your pod specification sets resources.limits.cpu: \u0026quot;2000m\u0026quot; (2 cores), Go still attempts to schedule goroutines across 64 concurrent OS threads. When traffic arrives, 64 threads wake up simultaneously, burn through the 200ms container quota in 3 milliseconds, and the kernel freezes the container for the remaining 97 milliseconds of the period.\nThe Fix: Automatic Cgroup Detection Import uber-go/automaxprocs in your main.go:\npackage main import ( _ \u0026#34;go.uber.org/automaxprocs\u0026#34; \u0026#34;net/http\u0026#34; ) func main() { // automaxprocs parses /sys/fs/cgroup/cpu.max automatically // and overrides GOMAXPROCS to match the cgroup quota: // limits.cpu: 2000m -\u0026gt; GOMAXPROCS=2 http.ListenAndServe(\u0026#34;:8080\u0026#34;, nil) } Memory Accounting: Why Containers Die Silently Unlike CPU, memory cannot be throttled. When a container exceeds its memory ceiling, the Linux kernel has only one recourse: invoking the Out-Of-Memory (OOM) killer.\nThe Anatomy of Exit Code 137 When Kubernetes terminates a pod with Exit Code 137 (\\(128 + 9 = \\text{SIGKILL}\\)), it does not allow the process to flush logs, write error traces, or trigger graceful shutdown handlers.\nThe termination decision is governed by kernel cgroup memory accounting:\nAnonymous Memory (RSS): Heap allocations, stack frames, goroutine structures. This memory cannot be evicted to disk. Page Cache: In-memory cached file reads, shared library segments, stdout/stderr socket buffers. When cgroup memory usage nears the limit, the kernel first attempts to reclaim page cache memory. If page cache memory is depleted and anonymous allocations continue to climb, the kernel selects a process within the cgroup based on oom_score_adj and sends an immediate, uncatchable SIGKILL.\nThe Go Memory Trap Prior to Go 1.19, the Go garbage collector operated purely on heap growth heuristics (GOGC=100). The runtime would trigger a GC cycle only when the heap grew by 100% relative to the live heap remaining after the previous collection.\nIf your container memory limit was set to 1GB, and live heap was 600MB, the Go GC would not schedule a collection until heap reached 1.2GB. The Linux kernel stepped in at 1.0GB and killed the container while the Go runtime sat idle, waiting for its collection threshold.\nThe Staff Fix: GOMEMLIMIT Introduced in Go 1.19, GOMEMLIMIT sets a hard ceiling on the Go runtime\u0026rsquo;s total memory footprint:\nenv: - name: GOMEMLIMIT # Set to ~85% of container memory limit to leave headroom for OS threads and page cache value: \u0026#34;850MiB\u0026#34; - name: GOGC value: \u0026#34;off\u0026#34; # Optional: Rely purely on GOMEMLIMIT for maximum throughput resources: requests: memory: \u0026#34;1Gi\u0026#34; limits: memory: \u0026#34;1Gi\u0026#34; With GOMEMLIMIT=850MiB, the Go garbage collector runs aggressively as memory approaches 850MB, preventing heap spikes from breaching the 1GB container limit and eliminating silent OOM kills.\nTechnology Trade-Off Matrix Resource Pattern Primary Advantage Operational Risk Production Verdict Strict CPU Limits (requests == limits) Absolute cost isolation; prevents noisy neighbors from consuming spare node cores. Destroys p99 latency during concurrent bursts; leads to artificial over-provisioning. Batch processing, asynchronous workers, untrusted multi-tenant clusters. Burstable CPU Limits (limits \u0026gt; requests) Improved resource sharing under moderate traffic variations. CFS throttling still triggers during peak bursts; unpredictable tail latency. Internal non-critical services with relaxed latency SLOs. No CPU Limits (requests set, limits.cpu: null) Zero CFS throttling; minimal tail latency; pods burst freely into idle node capacity. Buggy code with runaway loops (for {}) can saturate all unreserved cores on the host node. Client-facing synchronous microservices with strict p99 SLOs. Unbounded Memory (limits.memory: null) Pods never encounter container OOM kills. Memory leaks trigger node-wide memory exhaustion, causing Kubelet to evict unrelated pods. Architectural malpractice in production. The Staff-Level Decision Framework Top-tier engineering organizations (including Google, Zalando, and Shopify) have largely abandoned hard CPU limits on synchronous serving tiers.\nHere is the production standard for latency-critical workloads:\n1. Abolish CPU Limits on Synchronous APIs Set deterministic CPU requests to guarantee node capacity and scheduling priority, but leave CPU limits unset:\nresources: requests: cpu: \u0026#34;2000m\u0026#34; # Guaranteed baseline allocation memory: \u0026#34;2Gi\u0026#34; limits: # cpu: null \u0026lt;-- Do NOT set CPU limits on latency-sensitive serving paths memory: \u0026#34;2Gi\u0026#34; # Strict memory limit is mandatory 2. Protect Host Worker Nodes Without CPU limits, how do you prevent a runaway container from starving the host OS?\nConfigure node-level allocations in Kubelet:\n--system-reserved: Guarantees dedicated CPU and memory for sshd, systemd, and journald. --kube-reserved: Guarantees dedicated resources for kubelet and containerd. If an application enters an infinite loop, it bursts across available worker cores, but can never starve Kubelet or system daemons. Kubernetes continues to report metrics and can safely evict or restart workloads.\n3. Use Linux CPU Shares for Contention Arbitration When all pods on a node burst simultaneously, the Linux kernel uses CPU shares (derived directly from resources.requests.cpu) to allocate compute proportionally:\nPod A (requests.cpu: 2000m) receives twice as many CPU cycles as Pod B (requests.cpu: 1000m). Compute allocation remains fair and mathematically bounded without triggering arbitrary 100ms CFS sleep freezes. ","permalink":"https://hanhpham.vercel.app/posts/kubernetes-cfs-cpu-throttling-and-resource-limits/","summary":"Setting hard CPU limits on latency-sensitive microservices destroys p99 response times while average utilization appears deceptively low. Here is how Linux CFS quotas, Go GOMAXPROCS, and GOMEMLIMIT actually behave under load.","title":"CFS Quota Throttling \u0026 Silent OOMKills: The Fallacy of Kubernetes CPU Limits"},{"content":"You trigger a rolling deployment in production. The deployment strategy is configured with maxSurge: 25% and maxUnavailable: 0.\nYour monitoring graphs should show a smooth line. Instead, Datadog alerts fire: your ingress controller emits a burst of 502 Bad Gateway and 504 Gateway Timeout errors, while client mobile applications receive connection reset exceptions (ECONNRESET).\nAn engineer inspects the pod specification, searches StackOverflow, and commits a quick patch:\nlifecycle: preStop: exec: command: [\u0026#34;/bin/sh\u0026#34;, \u0026#34;-c\u0026#34;, \u0026#34;sleep 5\u0026#34;] The 502 errors decrease during low-traffic testing. The pull request gets merged.\nSix months later, during a major flash sale or marketing campaign, the errors return with triple the volume. Rolling updates now take twenty minutes because every replica waits on arbitrary sleeps, autoscaling cannot scale down fast enough to release cloud compute, and long-lived client connections still get aborted.\nUsing preStop: sleep 5 is not zero-downtime architecture; it is an unprincipled heuristic masking a distributed control plane race condition.\nThe 30-Second Architecture: When a Pod terminates, Kubernetes runs two independent asynchronous operations in parallel: local container teardown (SIGTERM from the Kubelet) and distributed network deregistration (EndpointSlice updates to kube-proxy and ingress controllers across all worker nodes). A sleep 5 hook attempts to guess the duration of that distributed network update. True zero-downtime deployments require application-level connection draining: intercepting SIGTERM, signaling upstream proxies via Connection: close (HTTP/1.1) or GOAWAY (HTTP/2), completing active requests, and exiting only when queues are dry.\nThe Distributed Race Condition: Why 502s Happen To eliminate deployment errors, you must understand the exact sequence of events that occurs when Kubernetes terminates a Pod replica.\nKubernetes is a distributed system governed by eventually consistent controllers. Pod termination does not happen in a linear, synchronous sequence.\n[ API Server: Pod Marked Terminating ] | +----------------------------+----------------------------+ | | v (Local Path) v (Distributed Network Path) [ Kubelet on Node A ] [ EndpointSlice Controller ] | | v v [ Sends SIGTERM to Container ] [ Updates EndpointSlice Object ] | | v v [ Container Exits Immediately ] [ Ingress / kube-proxy on all Nodes ] | | v v [ Sockets Closed / RST Sent ] [ Removes Pod IP from iptables/eBPF ] When a rolling update creates a new replica and targets an old replica for removal, two independent control loops execute concurrently:\nPath A: The Local Node Teardown The Kubelet on the local node observes that the Pod status is set to Terminating. The Kubelet removes the Pod from its local readiness checks. If a preStop hook is defined, the Kubelet executes it. Once preStop finishes (or immediately if none is defined), the Kubelet sends a SIGTERM signal to process ID 1 inside the container. If the container process has not exited after terminationGracePeriodSeconds (default 30 seconds), the Kubelet issues SIGKILL to force termination. Path B: The Distributed Network Deregistration The EndpointSlice controller detects the Pod\u0026rsquo;s deletion timestamp. The controller updates the EndpointSlice API object to mark the Pod as unready. Every worker node running kube-proxy detects the API change via its informer loop. Each kube-proxy rewrites its local iptables chains, IPVS tables, or Cilium eBPF map entries to stop routing new Service traffic to the Pod\u0026rsquo;s IP. The Ingress Controller (Envoy, Traefik, or Nginx Ingress) receives the event and updates its upstream connection routing pool. The Race Window Path A (local node) typically completes in 10 to 50 milliseconds if your application exits cleanly on SIGTERM.\nPath B (distributed network update across a 50-node cluster) takes 1 to 3 seconds depending on API server load, etcd write latency, and network propagation.\nDuring that 1-to-3-second window, your ingress controller and other cluster microservices still consider the terminating Pod healthy. They route fresh HTTP requests to the Pod IP. If the application process already exited during Path A, the Linux kernel on the node receives packets for a non-existent socket and responds with an immediate TCP RST. The client sees a 502 error.\nWhy preStop: sleep 5 Fails at Scale The naive response is to delay Path A by inserting a sleep into the preStop hook:\nlifecycle: preStop: exec: command: [\u0026#34;/bin/sh\u0026#34;, \u0026#34;-c\u0026#34;, \u0026#34;sleep 5\u0026#34;] This holds Path A for five seconds, giving Path B time to update cluster network endpoints. While this stops immediate connection resets on trivial HTTP/1.1 traffic, it introduces four severe operational costs:\n1. Slow Rollouts and Frozen Autoscaling Every replica destroyed during a deployment adds five seconds of mandatory idle wait. On a deployment with 40 replicas and maxUnavailable: 10%, a rolling update takes several extra minutes. When the Horizontal Pod Autoscaler (HPA) attempts to scale down unneeded compute after a traffic spike, nodes cannot be drained quickly, wasting cloud spend.\n2. The HTTP Keep-Alive Trap Modern HTTP clients and ingress proxies use persistent HTTP keep-alive connections. Envoy or Nginx maintains open TCP sockets to upstream pods to avoid three-way handshake overhead on every request.\nA sleep 5 hook does nothing to inform the ingress proxy that the connection should close. If a client sends a request at second 4.9, the application receives it right as the sleep expires and SIGTERM kills the process mid-stream.\n3. HTTP/2 and gRPC Stream Invalidation In HTTP/2 and gRPC architectures, hundreds of logical requests multiplex over a single persistent TCP connection. Abruptly terminating the underlying socket causes widespread stream failures across client services.\nTechnology Trade-Off Matrix Strategy Implementation Cost Cluster Impact HTTP/2 \u0026amp; gRPC Safety Production Recommendation No Lifecycle Hooks (Default) Zero Immediate 502 errors during every deployment Broken Dangerous in production preStop: sleep 5 Low (YAML edit) Masks race condition; adds 5s delay per replica teardown; ignores keep-alive pools Broken Prototypes / Non-critical batch jobs Readiness Probe Flipping Medium Pod drops out of endpoints before shutdown; polling interval delays drain Partial Acceptable fallback when code cannot be modified Application Connection Draining High (Requires code-level signal handling) Deterministic zero-downtime; exits as soon as in-flight requests finish Complete Mandatory for production microservices The Staff-Level Decision Framework: True Connection Draining To achieve true zero-downtime deployments without arbitrary sleeps, implement Coordinated Connection Draining inside your application runtime.\n[ 1. SIGTERM Received by App ] | v [ 2. Fail Local /healthz Endpoint ] (Drops out of local node checks immediately) | v [ 3. Signal Upstream Proxies to Stop ] (HTTP/1.1: Set \u0026#34;Connection: close\u0026#34; on responses) (HTTP/2 / gRPC: Emit GOAWAY frame to clients) | v [ 4. Drain Active In-Flight Requests ] (Process remaining queue; reject fresh keep-alives) | v [ 5. Close DB Connection Pools \u0026amp; Exit ] Go Implementation Pattern Here is how to implement deterministic connection draining in a Go HTTP service:\npackage main import ( \u0026#34;context\u0026#34; \u0026#34;net/http\u0026#34; \u0026#34;os\u0026#34; \u0026#34;os/signal\u0026#34; \u0026#34;syscall\u0026#34; \u0026#34;time\u0026#34; ) func main() { mux := http.NewServeMux() server := \u0026amp;http.Server{ Addr: \u0026#34;:8080\u0026#34;, Handler: mux, } // Intercept termination signals sigChan := make(chan os.Signal, 1) signal.Notify(sigChan, syscall.SIGINT, syscall.SIGTERM) go func() { if err := server.ListenAndServe(); err != nil \u0026amp;\u0026amp; err != http.ErrServerClosed { panic(err) } }() \u0026lt;-sigChan // 1. Create a timeout context bounded by Kubernetes terminationGracePeriodSeconds // Standard safety rule: timeout = terminationGracePeriodSeconds - 5s ctx, cancel := context.WithTimeout(context.Background(), 25*time.Second) defer cancel() // 2. server.Shutdown automatically: // - Stops accepting new TCP connections // - Sets \u0026#39;Connection: close\u0026#39; on open HTTP/1.1 connections // - Emits GOAWAY frames on HTTP/2 connections // - Waits for active in-flight requests to complete if err := server.Shutdown(ctx); err != nil { server.Close() } // 3. Close database pools, flush tracing spans, and exit cleanly } The Kubernetes Pod Configuration Once application connection draining is in place, configure the Kubernetes Deployment spec to match:\napiVersion: apps/v1 kind: Deployment metadata: name: billing-service spec: replicas: 10 strategy: type: RollingUpdate rollingUpdate: maxSurge: 25% maxUnavailable: 0 template: spec: # Must exceed application shutdown timeout by at least 5-10 seconds terminationGracePeriodSeconds: 35 containers: - name: app image: billing-service:v2.4.0 lifecycle: preStop: exec: # Small buffer (1-2s) to allow EndpointSlice propagation across worker nodes # before the application starts its internal drain command: [\u0026#34;/bin/sh\u0026#34;, \u0026#34;-c\u0026#34;, \u0026#34;sleep 2\u0026#34;] readinessProbe: httpGet: path: /healthz port: 8080 initialDelaySeconds: 3 periodSeconds: 2 failureThreshold: 2 The Rules to Remember Never set maxUnavailable \u0026gt; 0 on business-critical APIs: Always surge capacity before retiring old pods. terminationGracePeriodSeconds must be calculated mathematically: $$\\text{Grace Period} = \\text{Propagation Buffer (2s)} + \\text{Max Request Duration} + \\text{DB Cleanup Buffer (5s)}$$ Application runtimes must handle signals directly: If your container entrypoint is ENTRYPOINT [\u0026quot;/bin/sh\u0026quot;, \u0026quot;-c\u0026quot;, \u0026quot;./server\u0026quot;], bash absorbs SIGTERM and fails to forward it to your Go binary, forcing Kubernetes to kill your application with SIGKILL after 30 seconds. Always use the exec format: ENTRYPOINT [\u0026quot;./server\u0026quot;]. ","permalink":"https://hanhpham.vercel.app/posts/kubernetes-rolling-updates-and-connection-draining/","summary":"Slapping a five-second sleep into your container lifecycle masks a distributed control plane race condition. Here is how packet routing, EndpointSlice propagation, and application connection draining actually fit together.","title":"Zero-Downtime Kubernetes: Why preStop Sleep 5 Is an Architectural Anti-Pattern"},{"content":"A platform engineer wants to update a single DNS record or tweak an ingress annotation. They run terraform apply.\nTwelve minutes later, the plan finishes refreshing 840 cloud resources across three AWS regions. During those twelve minutes, the remote state lock on DynamoDB blocks every other engineer in the department from deploying. Even worse, an edge-case provider bug marks a shared transit gateway route as tainted, and the apply destroys network connectivity for six unrelated microservices.\nThe immediate reaction from leadership is predictable: \u0026ldquo;Our state file is too large. Break the monolith into micro-states.\u0026rdquo;\nTeams spend the next two quarters decomposing the infrastructure into fifty isolated directories. Plans drop from twelve minutes to eight seconds. Management celebrates.\nThen the integration tax arrives.\nThe 30-Second Architecture: Splitting state reduces the blast radius of a single apply, but moving state across directory boundaries shifts complexity into cross-stack data sharing. Using terraform_remote_state creates tight read-coupling, leaks root infrastructure secrets to application teams, and creates circular dependency deadlocks. Senior engineers split state to optimize plan execution speed. Staff engineers decouple state boundaries using typed, versioned parameter contracts (SSM Parameter Store, Vault) so stacks never inspect each other\u0026rsquo;s raw state files.\nThe Monolith State Problem: Why Big State Fails Monolithic state files fail because Terraform\u0026rsquo;s execution model scales poorly with resource volume.\nEvery terraform plan executes two phases:\nState Refresh: For each resource in terraform.tfstate, Terraform issues an API call to the cloud provider to fetch the live state. With 800 resources, you hit provider API rate limits and spend hundreds of seconds waiting on network round-trips. Graph Traversal: Terraform evaluates the Directed Acyclic Graph (DAG) of resource dependencies. A syntax error in a leaf resource can halt execution across the entire root graph. Acquiring state lock. This may take a few moments... Error: Error acquiring the state lock: ConditionalCheckFailedException: The conditional request failed Lock Info: ID: 84b1c8ee-8129-904c-6a7d-d6f15335c5eb Path: production-infrastructure/terraform.tfstate Operation: OperationTypeApply Who: alice@build-runner-04 Created: 2026-09-24 08:14:22.418291 +0000 UTC When four engineering squads share a single state file, lock contention paralyzes deployments. Engineers run terraform apply -lock=false out of frustration, causing concurrent state writes that corrupt the remote backend.\nThe Naive Fix: Micro-States \u0026amp; The Integration Tax To bypass lock contention, teams slice state into directories:\ninfrastructure/ ├── vpc/ ├── rds-postgres/ ├── redis-cluster/ ├── eks-cluster/ └── services/ ├── auth-service/ ├── payment-service/ └── order-service/ This fixes plan latency, but creates three systems-level failure modes:\n1. The terraform_remote_state Security Leak When auth-service needs the RDS endpoint and VPC subnet IDs, the standard tutorial recommendation is to read the upstream state:\n# The Naive Pattern: Reading upstream state directly data \u0026#34;terraform_remote_state\u0026#34; \u0026#34;rds\u0026#34; { backend = \u0026#34;s3\u0026#34; config = { bucket = \u0026#34;company-terraform-state\u0026#34; key = \u0026#34;rds-postgres/terraform.tfstate\u0026#34; region = \u0026#34;us-east-1\u0026#34; } } resource \u0026#34;aws_security_group_rule\u0026#34; \u0026#34;allow_auth\u0026#34; { security_group_id = data.terraform_remote_state.rds.outputs.db_security_group_id source_security_group_id = aws_security_group.auth.id } This introduces a severe security flaw: Terraform state files store all resource attributes in plain text, including sensitive outputs.\nTo read the security group ID, the IAM role executing auth-service must have s3:GetObject permissions on rds-postgres/terraform.tfstate. That state file also contains the raw master database password, KMS keys, and replication tokens. Slicing state into directories did not isolate privileges; it distributed root database credentials to every application runner.\n2. Cross-Stack Refactoring Deadlocks In a monolithic state, moving a resource between modules is a single code refactor.\nIn a decomposed multi-state setup, moving a resource across directory boundaries requires coordinated surgery:\nTarget the resource for deletion without destroying cloud infrastructure (terraform state rm aws_subnet.public_a). Navigate to the target directory. Import the live cloud resource into the target state (terraform import aws_subnet.public_a subnet-0123456789abcdef0). Update all downstream terraform_remote_state references across twelve application repos simultaneously. If an engineer makes a mistake during step 3, the target stack attempts to create the resource from scratch, throwing conflict errors from the cloud API.\n3. Circular Dependency Traps State A needs a resource from State B, while State B needs an output from State A (e.g. an EKS cluster needing an IAM role defined alongside an application worker, while the worker needs the EKS cluster OIDC issuer). Monolithic state handles this via topological sorting. Micro-states enter an unresolvable bootstrap deadlock.\nTechnology Trade-Off Matrix Dimension Monolithic State Micro-State Decomposition Decoupled Parameter Contracts Plan Latency 5 to 20 minutes (High lock contention) 10 to 30 seconds per directory 10 to 30 seconds per stack Blast Radius Catastrophic (Typo destroys shared core) Small (Confined to single component) Isolated (Strict boundary enforcement) Secret Isolation Zero (All state secrets accessible) Leaky (via terraform_remote_state) Complete (Fine-grained IAM on parameter keys) Refactor Cost Low (Single-pass code edits) High (Coordinated state mv calls) Medium (Versioned contract migrations) Tooling Sprawl Zero (Vanilla CLI) High (Terragrunt or wrapper shell scripts) Low (Native cloud data sources) Drift Visibility High (Single command detects drift) Low (Requires iterating 50 repositories) High (Scheduled CI drift sweeps) The Staff-Level Decision Framework Instead of coupling stacks via raw state files, structure infrastructure around Contract-Driven Decoupling.\n+-----------------------------+ | Foundation Tier (VPC / IAM) | +-----------------------------+ | v (Publishes typed contract IDs) +-------------------------------------------------------------+ | AWS SSM Parameter Store / HashiCorp Vault | | Key: /production/network/vpc_id | | Key: /production/network/private_subnets | +-------------------------------------------------------------+ ^ | (Reads typed strings; zero state file access) +-----------------------------+ | Application Tier (Services) | +-----------------------------+ 1. The Three-Tier Lifecycle Boundary Segment state strictly by rate of change and team blast radius:\nTier 1: Foundation (Quarterly changes): VPCs, transit gateways, Route53 public zones, IAM identity providers. Managed by core platform engineering. Tier 2: Platform (Monthly changes): EKS clusters, node groups, shared ingress controllers, cluster operators. Tier 3: Workloads (Daily/Hourly changes): Microservice deployments, database instances, SQS queues, S3 buckets. Managed by product squads. 2. Decouple Reads with Typed Parameter Contracts Ban data.terraform_remote_state in your linter. Upstream stacks publish their outputs as typed parameters in AWS SSM Parameter Store, Consul, or Vault:\n# Upstream (Foundation VPC Stack): Publish the contract resource \u0026#34;aws_ssm_parameter\u0026#34; \u0026#34;vpc_id\u0026#34; { name = \u0026#34;/production/network/vpc_id\u0026#34; type = \u0026#34;String\u0026#34; value = aws_vpc.main.id description = \u0026#34;Managed by terraform/foundation/vpc. Do not edit manually.\u0026#34; } Downstream application stacks consume the parameter using native cloud data sources:\n# Downstream (Application Stack): Read the contract data \u0026#34;aws_ssm_parameter\u0026#34; \u0026#34;vpc_id\u0026#34; { name = \u0026#34;/production/network/vpc_id\u0026#34; } resource \u0026#34;aws_security_group\u0026#34; \u0026#34;app\u0026#34; { vpc_id = data.aws_ssm_parameter.vpc_id.value } Why This Wins Least-Privilege Security: Application engineers only need read access to /production/network/* in SSM. They never get read permissions on the Foundation state file containing KMS keys or root credentials. Independent State Evolution: Upstream stacks can refactor their internal modules, rename resources, or switch from Terraform to OpenTofu without altering the SSM contract. Downstream stacks never experience breaking changes. Zero Circular Deadlocks: Inter-stack dependencies are resolved through standard cloud primitives rather than state file parsing. ","permalink":"https://hanhpham.vercel.app/posts/terraform-blast-radius-and-state-decomposition/","summary":"Splitting a monolithic terraform.tfstate into fifty micro-states speeds up plans, but introduces cross-stack read locks, secret leakage, and brittle dependencies. Here is how to structure state boundaries without the integration tax.","title":"Terraform State Decomposition: Blast Radius vs The Integration Tax"},{"content":"You want a binary decision in an automated workflow: should this customer support ticket escalate to engineering, or does this generated SQL query attempt an injection attack?\nYou construct a prompt with strict JSON schema instructions, add system constraints, and call a frontier model. Then you wait three seconds.\nThe model loads hundreds of context tokens, spins up its attention heads, and starts decoding. Because it is an autoregressive language model, it cannot simply return a boolean flag. It outputs text token by token:\nSure! Based on the provided database context, I have analyzed the user input... Even with structured outputs or JSON mode enabled, you watch the model run twenty sequential forward passes through seventy billion parameters just to emit twenty characters of syntax:\n{\u0026#34;is_escalation\u0026#34;: true, \u0026#34;confidence\u0026#34;: 0.98} Then your backend parser crashes because the model wrapped the JSON in markdown code fences, hallucinated an extra key, or dropped a closing bracket when hitting a token limit. You write regex fallbacks, wrap the call in retry loops, and pay twenty times what the computation was worth.\nUsing conversational language generation as the control plane for software logic is an architectural mismatch.\nThe 30-Second Architecture: Generative LLMs are sequential token predictors optimized to please human raters through Reinforcement Learning from Human Feedback (RLHF). For software control flow, natural language strings are an expensive, fragile anti-pattern. TypeSafe\u0026rsquo;s Jev replaces the autoregressive decoding loop with a non-autoregressive \u0026ldquo;System 1\u0026rdquo; engine: it encodes program state once, evaluates multiple typed questions (Bool, Choice, Score) across parallel decision heads in 70ms to 250ms, and trains on empirical accuracy via Reinforcement Learning from Calibrated Decisions (RLCD) rather than conversational vibes.\nTraditional autoregressive decoding versus parallel typed evaluation.\nThe Hidden Tax of Strings as Control Flow Every time software engineers integrate a generative LLM into a backend service, they pay an invisible three-part tax: latency, cost, and fragility.\n1. The Sequential Decoding Loop (\\(O(N)\\) Forward Passes) Generative transformers predict one token at a time. Each generated token requires reading all model weights from High Bandwidth Memory (HBM) into compute cores:\n$$\\text{Latency} = \\text{Time to First Token (TTFT)} + (N_{\\text{tokens}} \\times \\text{Inter-Token Latency})$$Even if your desired answer is a single enum value like \u0026quot;billing\u0026quot;, an autoregressive model in JSON mode still decodes roughly 15 to 30 syntax tokens sequentially ({, \u0026quot;, c, a, t, e, g, o, r, y, \u0026quot;, :, ...).\nAt 25ms per token on typical hosted infrastructure, generating twenty syntax tokens consumes 500ms of pure decoding time, entirely separate from the initial context prefill. When your backend handles thousands of automated triage requests per minute, that sequential loop creates massive bottlenecks.\n2. The String Parsing Failure Mode Software control flow requires rigid, deterministic types. LLMs emit variable-length unicode byte streams. The translation between those worlds is notoriously brittle:\npydantic_core._pydantic_core.ValidationError: 1 validation error for TicketRoutingDecision category Input should be \u0026#39;billing\u0026#39;, \u0026#39;technical\u0026#39;, or \u0026#39;sales\u0026#39;, got \u0026#39;Technical Support (Urgent)\u0026#39; [type=literal_error, input_value=\u0026#39;Technical Support (Urgent)\u0026#39;, input_type=str] To survive production, engineering teams build elaborate defensive wrappers:\nPrompting rules that threaten the model with penalties if it emits markdown fences. Grammar-constrained sampling engines (Outlines, Guidance) that force the logits to adhere to context-free grammars, adding CPU orchestration overhead. Multi-step retry loops that re-prompt the model when Pydantic parsing fails, multiplying latency and token bills. You are burning compute to force a text writer to act as an AST parser.\nThe Ex-OpenAI Thesis: Why RLHF Broke Decision Making The origin of Jev makes the architectural critique compelling.\nTypeSafe AI was founded by Diogo Almeida, a former OpenAI researcher, co-author of the original InstructGPT paper, and co-inventor of Reinforcement Learning from Human Feedback (RLHF), the alignment technique that enabled ChatGPT and GPT-4.\nAlmeida spent years building the mechanisms that taught language models to talk to people. Then he left to build an AI model that does not talk at all.\nHis core observation was straightforward: RLHF optimized models for conversational human satisfaction, which directly damaged their utility as software decision engines.\nWhen human raters grade AI responses during RLHF training, they systematically reward specific behaviors:\nPolite Verbosity: Longer, elaborately reasoned answers consistently receive higher ratings than terse, direct answers. False Confidence: Humans penalize models that say \u0026ldquo;I am uncertain\u0026rdquo; and reward answers delivered with unearned authority. Sycophancy: Models learn to agree with the user\u0026rsquo;s implicit premises rather than report contradictory ground truth. For a customer-facing chatbot, those behaviors create an engaging conversational persona. For software automation, they are toxic.\nWhen a routing engine asks an RLHF-trained model whether a transaction is fraudulent, the model\u0026rsquo;s reported probability is statistically uncalibrated. A model that assigns a 99% probability to an output might only be right 75% of the time. The numbers represent rhetorical confidence, not empirical probability.\nInside Jev: Non-Autoregressive System 1 Architecture To fix this, TypeSafe engineered Jev around Daniel Kahneman\u0026rsquo;s cognitive framework:\nSystem 2 (Slow, Deliberative): Frontier reasoning models (OpenAI o3, DeepSeek R1, Gemini 2.0 Thinking) that generate thousands of internal reasoning tokens to solve novel mathematical proofs or architect multi-file migrations. System 1 (Fast, Instinctive): Automated pattern recognition, classification, routing, and policy checks executed in milliseconds. Jev is built strictly as a System 1 model. It deletes natural language generation entirely.\nTraditional LLM: Context + Prompt --\u0026gt; Sequential Decoder --\u0026gt; \u0026#34;Here is the JSON: { ... }\u0026#34; --\u0026gt; Parser Error TypeSafe Jev: Context State --\u0026gt; Parallel Evaluation Heads --\u0026gt; { is_escalation: true (p=0.96) } 1. Encode Once, Evaluate Parallel Heads Instead of entering an iterative token generation loop, Jev takes two inputs:\nState Context: A shared string or structured data payload (code diff, support ticket, telemetry log, user query). Typed Questions: A dictionary of typed queries specifying the exact decision primitives required. The model encodes the state context in a single forward pass. Then, specialized classification and scoring heads evaluate all questions concurrently across the shared representation.\nBecause no text tokens are decoded, inference drops to 70ms to 250ms. There are no output tokens to bill or wait for.\n2. Supported Primitives Jev limits its output to typed mathematical primitives:\nBool: A boolean judgment returning true or false alongside calibrated confidence: \\(P(\\text{true})\\). Choice: Categorical selection among a predefined list of string labels, returning normalized probability distributions across all choices. Score: A continuous numeric value scaled against a defined rubric with confidence bounds. # Conceptual TypeSafe API Payload decision = await client.judge( state=user_ticket_payload, questions={ \u0026#34;needs_escalation\u0026#34;: Bool(criteria=\u0026#34;Customer account is Enterprise tier and service is completely degraded.\u0026#34;), \u0026#34;primary_category\u0026#34;: Choice(options=[\u0026#34;billing\u0026#34;, \u0026#34;outage\u0026#34;, \u0026#34;security\u0026#34;, \u0026#34;general_inquiry\u0026#34;]), \u0026#34;frustration_score\u0026#34;: Score(min_val=1, max_val=5, criteria=\u0026#34;Degree of customer distress in text.\u0026#34;) } ) # Returns native typed primitives directly if decision[\u0026#34;needs_escalation\u0026#34;].value and decision[\u0026#34;needs_escalation\u0026#34;].confidence \u0026gt; 0.90: queue.dispatch_pagerduty(ticket_id) No regex. No JSON validation. No markdown stripping.\nThe Meme: \u0026ldquo;My Name Is Jev\u0026rdquo; The departure from conversational chatbots leads to an immediate developer reaction: what happens when you ask Jev to write a poem, generate an essay, or talk about its feelings?\nIt cannot. It literally lacks an autoregressive language decoder.\nWhen developers attempt to use a non-autoregressive decision model as a conversational chatbot.\nIn the 2014 comedy 22 Jump Street, Channing Tatum attempts to infiltrate a high-stakes meeting with Mexican cartel leaders by adopting a terrible accent and repeating a single phrase: \u0026ldquo;My name is Jeff.\u0026rdquo; When pressed for details or complex conversation, the facade instantly breaks down.\nJev has the same relationship with natural language generation. If you feed it a writing prompt, it has no mechanism to output sentences. It does not converse; it evaluates state against typed predicates and returns probabilities.\nRLCD vs. RLHF: Calibrating for Truth The architectural engine that makes Jev useful is not merely speed; it is Reinforcement Learning from Calibrated Decisions (RLCD).\nIn standard RLHF, the reward model trains on human preference pairs: $$\\text{Reward} = f(\\text{Human A prefers Response 1 over Response 2})$$In RLCD, the model trains against ground-truth outcomes evaluated on scoring rules like the Brier score or binary log-loss:\n$$\\text{BS} = \\frac{1}{N} \\sum_{t=1}^{N} (f_t - o_t)^2$$Where \\(f_t\\) is the model\u0026rsquo;s forecasted probability and \\(o_t \\in \\{0, 1\\}\\) is the actual empirical outcome.\nMinimizing the Brier score forces the model to achieve probabilistic calibration: when Jev assigns an 80% confidence score across 1,000 different decisions, exactly 800 of those decisions must be empirically correct.\nWhy Calibration Enables Dual-Speed Architectures Uncalibrated probabilities are useless for automated pipelines. If an LLM claims it is \u0026ldquo;95% confident\u0026rdquo; but fails 30% of the time, you cannot trust it to run autonomous actions.\nWhen probabilities are statistically calibrated, software architects can establish mathematical risk thresholds:\n+-----------------------------+ | Incoming Request / Event | +-----------------------------+ | v +-----------------------------+ | Jev System 1 Model (80ms) | +-----------------------------+ | +--------------------+--------------------+ | | Confidence \u0026gt;= 0.95 Confidence \u0026lt; 0.95 | | v v +-------------------------------+ +-------------------------------+ | Automated Fast Execution | | Route to System 2 Model | | (Zero human/LLM delay) | | (Frontier LLM / o3 / Human) | +-------------------------------+ +-------------------------------+ High Confidence (\\(P \\ge 0.95\\)): Auto-execute the branch immediately. The task completes in under 100 milliseconds with zero human intervention. Medium/Low Confidence (\\(P \u003c 0.95\\)): Route the difficult edge case to an expensive System 2 reasoning model (o3-mini, Sonnet 3.7) or an on-call engineer. This dual-speed topology reduces frontier model API consumption by 80% to 90% while keeping end-to-end pipeline latency under 100ms for the vast majority of requests.\nArchitecture Comparison Feature Traditional Generative LLM TypeSafe Jev (System 1) Output Type Unstructured Unicode string / JSON Native typed primitives (Bool, Choice, Score) Execution Loop Autoregressive token-by-token decoding Single forward pass across parallel heads Latency Profile 1,500ms to 6,000ms 70ms to 250ms Output Billing Metered per output token Zero output token overhead Failure Mode Hallucinations, schema errors, lazy stubs Classification error (bounded by typed enum) Probabilistic Integrity Uncalibrated (RLHF sycophancy bias) Statistically calibrated (RLCD Brier loss) Ideal Workload Writing code, creative synthesis, open research Policy gating, security filtering, routing, triage Where Decision Models Fail Decision models are not a replacement for general intelligence. Understanding where the boundary lies prevents architectural misuse:\nZero Open-Ended Generation: If your feature requires drafting an email response, summarizing a transcript, or generating a code refactor, Jev cannot do it. You still need an autoregressive model. Dependent Sequential Logic: If Question B depends on the intermediate analytical result of Question A, evaluating them in a single parallel pass will fail. Complex multi-step reasoning requires either a System 2 thinking chain or chaining discrete decision calls. Rigid Schema Setup: You must know the questions and categorical choices at compile time. Jev cannot invent new categories on the fly. Treating language models as monolithic black boxes that must handle everything from UI conversation down to database routing is a historical accident of how LLMs entered software.\nText generation belongs at the human interface. Inside software architecture, strings are overhead; typed decisions are what matter.\nReferences \u0026amp; Further Reading Fireship: An ex-OpenAI researcher just deleted language from the LLM (Video teardown of TypeSafe AI and the Jev architecture). TypeSafe AI Official Site (System 1 models and non-autoregressive decision infrastructure). Training language models to follow instructions with human feedback (InstructGPT) (Ouyang, Almeida, et al., 2022: the foundational RLHF paper). Brier Score and Probability Calibration (Verification of probabilistic accuracy in decision models). Know Your Meme: My Name Is Jeff (Cultural origin of the 22 Jump Street quote). ","permalink":"https://hanhpham.vercel.app/posts/typesafe-jev-decision-models/","summary":"When you force a 70-billion-parameter language model to output \u0026rsquo;true\u0026rsquo; or \u0026lsquo;false\u0026rsquo;, you pay for sequential token decoding, uncalibrated probabilities, and broken JSON. Here is why TypeSafe stripped natural language out of the loop.","title":"Deleting Language from the LLM: Inside TypeSafe's Jev and Decision Models"},{"content":"You ask an assistant to refactor an HTTP handler, and it replaces your 500-line controller with a 40-line stub containing // ... rest of code remains the same. Or you ask it to fix a database query, and it confidently writes a solution using an ORM method deprecated three major versions ago. Or you run a multi-turn session across ten files and watch a $20 flat-rate subscription hit an opaque rate limit, while an API session burns thirty dollars in un-cached tokens. The default reaction is to blame the model. People go to social media, declare that the model or GPT got dumber over the weekend, and jump to a different editor. Almost every time, the model was not the problem. The failure happened two layers above it.\nAn AI coding setup is not a single product. It is a four-tier infrastructure stack. Conflating those tiers is why developers get burned by tools they do not understand.\nThe 30-Second Diagnosis: When an assistant breaks, 80% of failures originate in Layer 3 (the harness using vector RAG instead of LSP, or truncating patches) and Layer 2 (un-cached prefixes causing turn latency to spike from 1s to 15s). The model weights at Layer 1 are almost never the culprit.\nThe four-layer hierarchy: Surface, Harness, Gateway, and Base Model.\nLayer 1: The Model (Weights \u0026amp; Cognitive Specialization) At the bottom of the stack are the raw neural weights: Sonnet 3.7, OpenAI o3-mini, DeepSeek R1, Gemini 2.0 Flash.\nModels are not interchangeable commodities that differ only on a benchmark leaderboard. They have distinct cognitive profiles:\nMulti-File Spatial Reasoning: Sonnet 3.7 excels at understanding sprawling repository structure, matching existing code conventions, and respecting architectural boundaries. Algorithmic Constraint Solvers: OpenAI o3-mini and o1 excel at pure logic puzzles, self-contained algorithms, and competitive programming problems with rigid mathematical invariants. Cost-Weighted Inference: DeepSeek R1 provides deep reasoning traces at roughly a tenth of the API cost of Western frontier models, making brute-force exploratory passes affordable. The model is responsible for exactly one job: predicting the next token given a context window. It does not read your files, it does not run your tests, and it does not write to your disk. When an assistant hallucinates a function signature, it is rarely because the weights cannot reason; it is because the layers above failed to supply the interface definition in the prompt.\nLayer 2: The Gateway \u0026amp; Inference Provider (The Economics of Caching) The model is hosted behind a gateway: direct vendor APIs (OpenAI, Google, Mistral), proxy aggregators (OpenRouter), enterprise hyperscalers (AWS Bedrock, Azure AI), or local engines (vLLM, Ollama). Developers treat this layer as an invisible pipe. It is actually the difference between an assistant you can afford to run and one that bankrupts you.\nThe load-bearing feature at Layer 2 is prompt caching.\nConsider a realistic agent session:\nSystem prompt, tool definitions, rules: ~8,000 tokens Repository map, project instructions: ~12,000 tokens Five open source files and type definitions: ~30,000 tokens Total base context: 50,000 tokens If your assistant runs a 15-step agentic loop to debug a failing test, it calls the API 15 times.\nWithout prompt caching, you pay for all 50,000 input tokens on every turn: $$15 \\times 50,000 = 750,000 \\text{ input tokens}$$ On a frontier reasoning model ($3 per million input tokens), that single debugging task costs $2.25, and every turn takes 12 to 18 seconds of time-to-first-token latency while the provider processes the prefix.\nWith prompt caching (where providers like OpenAI and frontier gateways offer a 90% discount on cache hits):\nTurns 2 through 15 hit the cache: 50,000 tokens at a 90% discount ($0.015 per turn) Total input cost: $0.36 (an 84% cost drop) Time-to-first-token drops from 15 seconds to under 1.5 seconds. Prompt caching is not a minor cost optimization. It is the architectural prerequisite for multi-turn autonomous coding. If your provider drops cache keys, rate-limits prefix writes, or does not support prompt caching, agentic coding becomes unusable.\nLayer 3: The Harness \u0026amp; Agent Loop (Where 80% of Failures Happen) The harness is the orchestrator: Aider, OpenCode, Cline, Devin, and terminal coding harnesses. This is the engine room of the stack, and it is where almost all assistant failures originate. A harness owns three responsibilities: context retrieval, patch application, and the tool execution loop.\n1. The Vector Search Trap (RAG vs LSP) Early coding assistants relied heavily on vector embeddings for codebase context. You type a prompt, the harness converts it to a vector, searches a vector database of code chunks, and pastes the top five semantic matches into the prompt.\nFor software engineering, pure vector RAG is fundamentally broken. Embeddings understand semantic similarity, not software architecture:\nIf you ask to modify HandleUserLogin, vector search grabs five files that contain the word \u0026ldquo;login\u0026rdquo; (comments, documentation, auth configs). It completely misses the User struct definition, the database interface, and the middleware call site located in completely different directories. Modern harnesses have largely abandoned pure vector search in favor of deterministic tools:\nripgrep and file discovery: Fast regex sweeps across filenames and content. Language Server Protocol (LSP): Querying the actual compiler/language server for goToDefinition, findReferences, and typeDefinition. This follows the real call graph instead of guessing based on word proximity. AST Parsing (tree-sitter): Structural syntax parsing that extracts class skeletons and function signatures without burning context on implementation bodies. 2. Patch Generation Strategies Once the model decides what code to write, how does the harness apply it to your disk?\nThere are three main strategies:\nStrategy How It Works The Failure Mode Whole-File Rewrite The model outputs the entire file from line 1 to the end. The Lazy Stub: On files longer than 200 lines, the model runs out of output tokens or patience and emits // ... existing code ..., destroying your implementation. Search and Replace The model outputs a search block and a replace block. The Fragility Trap: If indentation shifts by two spaces, or if the search block matches three different places in the file, the patch fails or corrupts code. Line-Anchored Patching The harness copies line tags from latest reads and applies anchored surgical edits. Zero hallucination of unchanged lines; lowest token output; fails safely if stale lines are touched. When an assistant corrupts your file, it was almost never a reasoning failure in the model. It was a patch application failure in the harness.\n3. The Tool Execution Loop \u0026amp; Sandboxing An assistant that cannot execute tools is flying blind. A capable harness runs a closed loop:\nPropose edit. Apply patch to file. Run project linter or compiler (cargo check, go test ./..., npm run build). Read compiler errors back into the context window. Self-correct before yielding to the user. Without this loop, the user acts as the compiler, manually copying errors back and forth between terminal and chat box.\n4. What real benchmarks actually measure: scaffolds over weights Early AI coding benchmarks (like HumanEval or MBPP) tested single functions in isolation: the model was given a docstring and asked to generate a Python function.\nModern evaluations look entirely different because real software engineering is an integration test:\nAider\u0026rsquo;s Polyglot \u0026amp; Refactoring Benchmarks: The Polyglot benchmark tests 225 problems across six languages (C++, Go, Java, JS, Python, Rust) in isolated containers with test feedback loops. More importantly, Aider\u0026rsquo;s Refactor benchmark asks models to refactor 89 massive methods, specifically testing whether the harness and model can output dense code without truncating lines, dropping syntax, or taking lazy shortcuts. SWE-bench Verified \u0026amp; SWE-bench Pro: Evaluating models on hundreds of real GitHub issues. The most telling finding from the SWE-bench leaderboard is scaffold sensitivity: taking the exact same base model and pairing it with an optimized execution harness (such as mini-SWE-agent with structured bash tools and iterative retries) swings issue resolution rates by 20 to 30 percentage points. When frontier models reach 70% to 80%+ on SWE-bench Verified, the harness is doing half the heavy lifting.\nLayer 4: The Interaction Surface (IDEs, Extensions, and CLI) The top layer is what you actually look at:\nForked AI IDEs: Cursor, Windsurf. Terminal \u0026amp; CLI Native: Aider, OpenCode, and dedicated terminal coding harnesses. Each surface optimizes for an entirely different phase of engineering work. GUI IDEs Win at Micro-Edits Cursor and Windsurf provide a polished experience for single-file, interactive development:\nShadow Workspaces: Cursor runs a background language server in a hidden workspace to check compiler diagnostics while the model streams text. Speculative Tab-Completion: Predicting where your cursor will jump next and pre-streaming ghost text. Inline Visual Diffs: Reviewing green and red diff hunks directly in the editor buffer with keyboard shortcuts (Ctrl + K). The downside: Cursor is a closed-source downstream fork of VS Code. You are permanently dependent on a third-party vendor keeping up with upstream VS Code security patches, extension marketplace quirks, and enterprise SSO integrations.\nTerminal-native harnesses (like Aider, OpenCode, or headless CLI agents) operate directly inside your shell:\nThey do not compete with your editor. You keep your personalized Neovim, Emacs, or VS Code setup intact. They excel at multi-file migrations, sweeping test suites, and autonomous background execution where visual GUI chrome only adds overhead. The $20 Flat-Rate Trap vs Pay-Per-Token Reality One of the biggest points of confusion among developers is the economic divide between flat-rate subscriptions and direct API access.\nThe Flat-Rate Tier (Copilot, Cursor Pro: $20/month) Providers cannot afford to give you unlimited raw frontier models for twenty dollars. Behind the scenes, flat-rate tiers rely on aggressive cost controls: hard rate limits, request queues during peak hours, smaller fallback models, and speculative routing. This is fine for interactive autocomplete and small edits. It falls apart during heavy multi-file architectural refactors that consume millions of context tokens. Pay-Per-Token APIs (Aider, OpenCode, CLI Agents) There are no opaque rate limits or hidden model downgrades during peak hours. With Layer 2 prompt caching, an intensive 20-step refactor across a codebase might cost $0.80 to $1.50. You pay for what you actually execute, with full visibility into input, output, and cache hit metrics. Diagnosing the Failure The next time an AI assistant produces garbage, resist the urge to declare that \u0026ldquo;the model is broken\u0026rdquo;. Walk down the stack:\nDid the Surface fail? Did the editor drop the cursor context or truncate the inline diff? Did the Harness fail? Did it retrieve irrelevant vector chunks instead of symbol definitions? Did its patch format hallucinate code while rewriting a 400-line file? Did it skip running tests? Did the Gateway fail? Did a proxy drop your prompt cache, causing turn latency to blow up from 1 second to 15 seconds? Did the Model fail? Was the algorithmic reasoning fundamentally flawed despite clean context and precise tools? Eighty percent of the time, the failure lives in Layers 2 and 3. When you understand the stack, you stop cargo-culting tools and start building an environment that actually survives a production codebase.\n","permalink":"https://hanhpham.vercel.app/posts/the-four-layer-ai-coding-stack/","summary":"Most developer frustration with AI coding assistants comes from conflating the editor, the agent loop, the gateway, and the model. Here is how the four layers actually fit together, where failures originate, and how to build a setup that survives production work.","title":"The Four-Layer AI Coding Stack: Why Your Assistant Fails (and It's Not the Model)"},{"content":"Yank five lines from a 120-column tmux pane inside WSL2, paste them into a Python file or YAML manifest in your editor, and your linter immediately flags eighty trailing spaces. Copy a command from your shell history, paste it into bash, and the shell halts with command not found: \\u00a0... because your prompt injected an invisible non-breaking space. Copy a multi-line snippet back from Windows, and bare carriage returns mangle your terminal line breaks.\nStandard Linux clipboard utilities like xclip or wl-copy fail out of the box in headless WSL2 because there is no X11 or Wayland display server. The standard advice online is piping tmux selections directly to /mnt/c/WINDOWS/system32/clip.exe. That works for about ten minutes, until the trailing whitespace, broken prompt glyphs, and line ending mismatches turn daily development into death by a thousand paper cuts.\nA clean setup needs a bidirectional pipeline: sanitize text on the way out, normalize line endings on the way in, and complete in single-digit milliseconds so mouse dragging never stutters.\nThe 30-Second Setup: If your terminal clipboard is broken right now and you just want the working code: grab tmux-copy-wsl and tmux-paste-wsl, drop them into ~/.local/bin/, make them executable, and append the tmux bindings to ~/.tmux.conf. If you want to understand why standard pipes corrupt prompt glyphs, freeze server loops, and leak clipboard history, read the breakdown below.\nThe three invisible characters breaking your pastes Before writing any configuration, here is the exact garbage raw terminal copying produces when you drag a cursor across a tmux pane:\nFixed-width column padding: tmux allocates fixed columns for each pane. If your pane is 120 columns wide and you yank a 25-character command, tmux hands your clipboard 95 trailing spaces. Paste that into YAML, Python, or markdown, and you get immediate indentation bugs or noisy git diffs. Prompt glyphs and non-breaking spaces: Modern prompt engines (omp, Starship, Powerline) use non-breaking spaces (\\u00a0, \\u202f, \\u2007) to keep prompt segments contiguous, along with zero-width characters (\\u200b through \\u200d, \\ufeff) for icon alignment. If you copy a command from your terminal scrollback, those code points come with it. Bash halts with command not found: \\u00a0... or fails with mysterious token errors. Carriage returns (\\r\\n): Text copied on the Windows host carries CRLF line breaks. In a Linux shell, bare carriage returns cause cursor jumps, broken heredocs, and script parse failures. Bidirectional data flow: low-latency Perl sanitization on copy, bracketed paste and carriage return normalization on paste.\nSanitizing on copy: why Perl beats Python and sed The copy helper sits directly inside the mouse drag and vi-yank loop. If the sanitizer takes 45 milliseconds to start, every single mouse release stutters.\nI tried Python first. It took 35 to 50 milliseconds just to boot the runtime and import standard libraries. On every mouse release, that latency was clearly perceptible: it felt like dragging the cursor through wet cement. Standard POSIX sed and awk start in under 2 milliseconds, but POSIX sed lacks consistent multi-byte Unicode range matching across different locales, frequently mangling UTF-8 sequences or failing on multi-byte code point ranges like \\x{200b} through \\x{200d}.\nPerl is pre-installed on virtually every Linux distribution (including stock Ubuntu WSL2 images), boots in roughly 2 milliseconds, and handles UTF-8 natively with -CSD.\nSave this script as ~/.local/bin/tmux-copy-wsl:\n#!/usr/bin/env bash # tmux-copy-wsl: Sanitize text yanked from tmux/omp and copy to Windows clipboard + tmux buffer set -eo pipefail # Prioritize canonical system path to avoid PATH hijacking from weak directories if [ -x \u0026#34;/mnt/c/WINDOWS/system32/clip.exe\u0026#34; ]; then CLIP_CMD=\u0026#34;/mnt/c/WINDOWS/system32/clip.exe\u0026#34; else CLIP_CMD=\u0026#34;$(command -v clip.exe 2\u0026gt;/dev/null || true)\u0026#34; fi # Sanitize text from stdin via Perl (UTF-8 native, fast startup): # 1. Normalize line endings (\\r\\n and \\r to Unix \\n) # 2. Strip zero-width characters (\\u200b, \\u200c, \\u200d, \\ufeff) # 3. Convert non-breaking spaces (\\u00a0, \\u202f, \\u2007) to standard spaces # 4. Strip redundant trailing spaces and tabs from each line (fixes tmux pane padding) # 5. Normalize trailing blank lines to a single newline # # Note on command substitution: bash $(...) unconditionally strips trailing newlines. # Appending a sentinel character (\u0026#39;x\u0026#39;) and stripping it afterwards preserves exact newlines. cleaned=$(perl -CSD -0777 -pe \u0026#39; s/\\r\\n/\\n/g; s/\\r/\\n/g; s/[\\x{200b}-\\x{200d}\\x{feff}]//g; s/[\\x{00a0}\\x{202f}\\x{2007}]/ /g; s/[ \\t]+$//mg; s/\\n+$/\\n/; \u0026#39; 2\u0026gt;/dev/null || cat; printf x) cleaned=\u0026#34;${cleaned%x}\u0026#34; if [ -n \u0026#34;$cleaned\u0026#34; ]; then printf \u0026#39;%s\u0026#39; \u0026#34;$cleaned\u0026#34; | tmux load-buffer - 2\u0026gt;/dev/null || true if [ -n \u0026#34;$CLIP_CMD\u0026#34; ] \u0026amp;\u0026amp; [ -x \u0026#34;$CLIP_CMD\u0026#34; ]; then printf \u0026#39;%s\u0026#39; \u0026#34;$cleaned\u0026#34; | \u0026#34;$CLIP_CMD\u0026#34; 2\u0026gt;/dev/null || true fi fi Make it executable:\nchmod +x ~/.local/bin/tmux-copy-wsl A subtle footgun with bash command substitution: var=$(...) unconditionally strips all trailing newlines. If you yank a whole line with TripleClick or copy multiple lines, Perl normalizes the trailing newline, and then bash immediately deletes it. The fix is the classic sentinel character idiom: append an x inside the subshell, then strip it with ${cleaned%x}. Because x is the final character, bash leaves every preceding newline intact.\nBreaking down the Perl flags and substitutions:\n-CSD: Tells Perl that standard input, standard output, and standard error streams are encoded in UTF-8. Without this flag, Perl treats multi-byte characters as raw byte sequences, corrupting multi-byte glyphs. -0777: Enables slurping mode so the entire input is ingested into memory as a single string, allowing multiline matching (/m) and cross-line replacements. s/[\\x{200b}-\\x{200d}\\x{feff}]//g: Deletes zero-width spaces, non-joiners, joiners, and byte order marks while keeping legitimate multi-byte characters (such as accented letters or Nerd Font icons) untouched. s/[\\x{00a0}\\x{202f}\\x{2007}]/ /g: Replaces non-breaking space variants with plain ASCII space (0x20). s/[ \\t]+$//mg: Strips trailing spaces and tabs before each newline, stripping tmux column padding without touching indentation at the beginning of lines. Pasting back from Windows: PowerShell and bracketed paste Reading the Windows clipboard from inside WSL2 requires querying the Win32 clipboard API through powershell.exe.\nThe naive one-liner most people use:\npowershell.exe -Command \u0026#34;Get-Clipboard\u0026#34; This breaks in three distinct ways:\nLine arrays: Without -Raw, PowerShell splits multi-line text into an array of strings, mangling formatting. Silent ASCII degradation: On Windows 10 and 11, Windows PowerShell 5.1 defaults $OutputEncoding to US-ASCII when output is redirected across a standard Linux pipe. Any non-ASCII character (Vietnamese accents, emoji, smart quotes) turns into a literal question mark (?). Server-wide freezes: In tmux, run-shell runs synchronously by default. If PowerShell hangs on host clipboard mutex contention or a virus scan, the entire tmux server stops processing keystrokes across all panes and windows. Save this hardened script as ~/.local/bin/tmux-paste-wsl:\n#!/usr/bin/env bash # tmux-paste-wsl: Read text from Windows clipboard and paste into active tmux pane set -eo pipefail # Prioritize canonical Windows PowerShell path if [ -x \u0026#34;/mnt/c/WINDOWS/System32/WindowsPowerShell/v1.0/powershell.exe\u0026#34; ]; then POWERSHELL_CMD=\u0026#34;/mnt/c/WINDOWS/System32/WindowsPowerShell/v1.0/powershell.exe\u0026#34; else POWERSHELL_CMD=\u0026#34;$(command -v powershell.exe 2\u0026gt;/dev/null || true)\u0026#34; fi # Query host clipboard with a hard timeout to prevent blocking the tmux server event loop. # 1. Force PowerShell OutputEncoding to UTF-8 to prevent non-ASCII characters from becoming \u0026#39;?\u0026#39;. # 2. Normalize CRLF and bare CR to Unix LF. # 3. Repeatedly strip bracketed paste delimiters (\\e[200~, \\e[201~, \\x9b200~, \\x9b201~) to neutralize nested injections. # 4. Strip \\e[?2004l sequences to prevent payloads from deactivating terminal bracketed paste mode. if [ -n \u0026#34;$POWERSHELL_CMD\u0026#34; ] \u0026amp;\u0026amp; [ -x \u0026#34;$POWERSHELL_CMD\u0026#34; ]; then timeout 2s \u0026#34;$POWERSHELL_CMD\u0026#34; -NoProfile -Command \u0026#34; [Console]::OutputEncoding = [System.Text.Encoding]::UTF8; Get-Clipboard -Raw \u0026#34; 2\u0026gt;/dev/null \\ | perl -pe \u0026#39; s/\\r\\n?/\\n/g; 1 while s/(?:\\x1b\\[|\\x9b|\\xc2\\x9b)20[01]~//g; s/(?:\\x1b\\[|\\x9b|\\xc2\\x9b)\\?2004[lh]//g; \u0026#39; \\ | tmux load-buffer - 2\u0026gt;/dev/null || true fi # Paste with bracketed paste (-p) explicitly into the originating pane tmux paste-buffer -p -t \u0026#34;${TMUX_PANE:-.}\u0026#34; 2\u0026gt;/dev/null || true Make it executable:\nchmod +x ~/.local/bin/tmux-paste-wsl Security note on bracketed paste: The -p flag on tmux paste-buffer -p is an essential terminal security control. It wraps pasted text in ANSI bracketed paste escape sequences (\\e[200~ and \\e[201~). If you paste a multi-line snippet or code with embedded newlines, bracketed paste prevents the shell from executing those lines as immediate shell commands before you can inspect them.\nWiring it into ~/.tmux.conf With the helper scripts in place, add the following section to your ~/.tmux.conf:\n# ============================================================================== # WSL2 \u0026amp; Windows Host Mouse Mode and Clipboard Integration # ============================================================================== # Enable mouse mode and OSC 52 terminal clipboard integration set -g mouse on set -s set-clipboard on setw -g mode-keys vi # Vi copy-mode keybindings for manual selection bind-key -T copy-mode-vi v send-keys -X begin-selection bind-key -T copy-mode-vi C-v send-keys -X rectangle-toggle # Yank with \u0026#39;y\u0026#39; or \u0026#39;Enter\u0026#39; using sanitized copy filter to Windows clipboard bind-key -T copy-mode-vi y send-keys -X copy-pipe-and-cancel \u0026#34;tmux-copy-wsl\u0026#34; unbind-key -T copy-mode-vi Enter bind-key -T copy-mode-vi Enter send-keys -X copy-pipe-and-cancel \u0026#34;tmux-copy-wsl\u0026#34; # Mouse drag selection copies immediately upon release and syncs to Windows clipboard bind-key -T copy-mode-vi MouseDragEnd1Pane send-keys -X copy-pipe-and-cancel \u0026#34;tmux-copy-wsl\u0026#34; bind-key -T copy-mode MouseDragEnd1Pane send-keys -X copy-pipe-and-cancel \u0026#34;tmux-copy-wsl\u0026#34; # Double-click (copy word) and Triple-click (copy line) in copy-mode bind-key -T copy-mode-vi DoubleClick1Pane select-pane \\; send-keys -X select-word \\; run-shell -d 0.3 \\; send-keys -X copy-pipe-and-cancel \u0026#34;tmux-copy-wsl\u0026#34; bind-key -T copy-mode-vi TripleClick1Pane select-pane \\; send-keys -X select-line \\; run-shell -d 0.3 \\; send-keys -X copy-pipe-and-cancel \u0026#34;tmux-copy-wsl\u0026#34; bind-key -T copy-mode DoubleClick1Pane select-pane \\; send-keys -X select-word \\; run-shell -d 0.3 \\; send-keys -X copy-pipe-and-cancel \u0026#34;tmux-copy-wsl\u0026#34; bind-key -T copy-mode TripleClick1Pane select-pane \\; send-keys -X select-line \\; run-shell -d 0.3 \\; send-keys -X copy-pipe-and-cancel \u0026#34;tmux-copy-wsl\u0026#34; # Double-click (word) and Triple-click (line) from a normal pane (starts copy-mode and yanks) bind-key -n DoubleClick1Pane select-pane -t = \\; if-shell -F \u0026#34;#{||:#{pane_in_mode},#{mouse_any_flag}}\u0026#34; \u0026#34;send-keys -M\u0026#34; \u0026#34;copy-mode -H ; send-keys -X select-word ; run-shell -d 0.3 ; send-keys -X copy-pipe-and-cancel tmux-copy-wsl\u0026#34; bind-key -n TripleClick1Pane select-pane -t = \\; if-shell -F \u0026#34;#{||:#{pane_in_mode},#{mouse_any_flag}}\u0026#34; \u0026#34;send-keys -M\u0026#34; \u0026#34;copy-mode -H ; send-keys -X select-line ; run-shell -d 0.3 ; send-keys -X copy-pipe-and-cancel tmux-copy-wsl\u0026#34; # Mouse paste: Middle-click (button 2) and Right-click (button 3) paste Windows clipboard bind-key -n MouseDown2Pane select-pane -t = \\; if-shell -F \u0026#34;#{||:#{pane_in_mode},#{mouse_any_flag}}\u0026#34; \u0026#34;send-keys -M\u0026#34; \u0026#34;run-shell -b tmux-paste-wsl\u0026#34; bind-key -n MouseDown3Pane select-pane -t = \\; if-shell -F \u0026#34;#{||:#{pane_in_mode},#{mouse_any_flag}}\u0026#34; \u0026#34;send-keys -M\u0026#34; \u0026#34;run-shell -b tmux-paste-wsl\u0026#34; # Mouse paste while inside copy-mode cancels copy-mode and pastes bind-key -T copy-mode-vi MouseDown2Pane send-keys -X cancel \\; run-shell -b tmux-paste-wsl bind-key -T copy-mode-vi MouseDown3Pane send-keys -X cancel \\; run-shell -b tmux-paste-wsl bind-key -T copy-mode MouseDown2Pane send-keys -X cancel \\; run-shell -b tmux-paste-wsl bind-key -T copy-mode MouseDown3Pane send-keys -X cancel \\; run-shell -b tmux-paste-wsl # Keyboard paste from Windows clipboard: Prefix + ] bind-key ] run-shell -b tmux-paste-wsl Reload tmux inside an active session:\ntmux source-file ~/.tmux.conf How the normal-mode click detection works The double and triple click bindings for normal panes use an interesting tmux construct:\nbind-key -n DoubleClick1Pane select-pane -t = \\; if-shell -F \u0026#34;#{||:#{pane_in_mode},#{mouse_any_flag}}\u0026#34; \u0026#34;send-keys -M\u0026#34; \u0026#34;copy-mode -H ; send-keys -X select-word ; run-shell -d 0.3 ; send-keys -X copy-pipe-and-cancel tmux-copy-wsl\u0026#34; Four moving parts make this work:\nselect-pane -t =: Activates the exact pane currently under the mouse pointer. if-shell -F \u0026quot;#{||:#{pane_in_mode},#{mouse_any_flag}}\u0026quot;: Tests whether the pane is already in a special mode or running an application that captures mouse events (such as less, vim, or htop). If so, send-keys -M forwards the mouse event directly to the running application without interfering. copy-mode -H: Enters copy mode with the -H flag to hide the top-right position indicator during rapid selection. run-shell -d 0.3: Introduces a 300-millisecond delay so tmux has time to visually highlight the selected word before copy-pipe-and-cancel captures the text and exits copy mode. Security considerations: injection, timeouts, and clipboard history Bridging an untrusted host clipboard to a shell running with user privileges requires defensive precautions:\n1. Bracketed paste escape sequence injection \u0026amp; nested delimiters Terminal emulators rely on bracketed paste delimiters (\\e[200~ and \\e[201~) to signal when input should be treated as literal text rather than typed commands. However, naive filtering presents two immediate evasion vectors:\nNested Delimiter Smuggling: A simple single-pass substitution (s/\\x1b\\[201~//g) is vulnerable to nesting. If an attacker crafts \\e[\\e[201~201~, stripping the inner token reconstructs \\e[201~ in the output stream. The sanitizer resolves this by wrapping substitution in a loop (1 while s/...//g). 8-Bit C1 CSI Controls: In addition to the standard two-byte 7-bit escape (\\e[ or \\x1b[), terminal emulators support single-byte 8-bit C1 CSI sequences (\\x9b or UTF-8 \\xc2\\x9b). An attacker using \\x9b201~ bypasses filters that only look for \\x1b[. The regex covers both variants. Mode Deactivation (\\e[?2004l): Terminals track bracketed paste status using DEC Private Mode 2004. An escape sequence containing \\e[?2004l instructs the terminal to turn off bracketed paste mode immediately. Stripping this sequence preserves terminal state. 2. PowerShell 5.1 default ASCII output encoding On Windows 10 and 11, the built-in Windows PowerShell 5.1 host defaults $OutputEncoding to US-ASCII when output is redirected to a standard pipe. Without intervention, any non-ASCII characters copied on Windows (such as Vietnamese characters, accented text, smart quotes, or emoji) are converted to literal question marks (?). Setting [Console]::OutputEncoding = [System.Text.Encoding]::UTF8; inside the command guarantees full Unicode fidelity across the boundary.\n3. Server-wide event loop lockups: synchronous vs background run-shell In tmux, run-shell without the -b flag runs synchronously, halting the entire tmux server event loop until the invoked script exits. Because tmux-paste-wsl crosses the virtualization boundary into Windows PowerShell, any host clipboard mutex contention, disk spike, or antivirus scan would freeze all panes, windows, and attached sessions.\nTwo safeguards eliminate this risk:\nKeybindings use run-shell -b, spawning the helper asynchronously so tmux continues handling user keystrokes. The helper wraps PowerShell in timeout 2s, ensuring a hung process is forcibly killed if the host does not respond within two seconds. 4. Child process bracketed paste support tmux paste-buffer -p only wraps pasted text in bracketed paste sequences if the application running inside the pane has requested bracketed paste mode (such as modern bash with readline, zsh, fish, vim, or python). If you paste into a legacy application or raw shell prompt that does not enable mode 2004, the text is emitted raw.\n5. WSL2 PATH precedence hijacking WSL2 automatically imports Windows system folders into Linux $PATH. If an unprivileged user or cloned project creates a weak directory earlier in $PATH containing a mock clip.exe or powershell.exe, resolving via bare command -v executes the local script instead of the Windows host binary.\nChecking /mnt/c/WINDOWS/system32/clip.exe and /mnt/c/WINDOWS/System32/WindowsPowerShell/v1.0/powershell.exe explicitly ensures only the genuine Windows system binaries are executed.\n6. Windows Clipboard History (Win + V) and cloud synchronization When you yank text into clip.exe, Windows adds that string to its clipboard. If you have Windows Clipboard History enabled, any sensitive strings (such as API keys, database passwords, or temporary tokens) are written to local host storage (%LOCALAPPDATA%\\Microsoft\\Windows\\Clipboard). If cloud clipboard sync is turned on, those credentials are automatically uploaded to Microsoft cloud services and synced across other devices logged into your account.\nRule of thumb: For sensitive credentials or private keys, hold Shift while dragging to bypass tmux mouse capture and copy locally inside the terminal emulator, or clear your Windows clipboard history immediately after yanking.\nDaily usage cheat sheet Action Input Behavior Mouse Drag Selection Left click and drag, release Sanitizes selection and copies to both tmux buffer and Windows clipboard Copy Single Word Double-click word Selects word, trims padding, and copies Copy Full Line Triple-click line Selects line, strips trailing column spaces, and copies Keyboard Yank Prefix + [ -\u0026gt; v -\u0026gt; move -\u0026gt; y Copies selection into Windows clipboard Mouse Paste (Windows style) Right-click in pane Pastes Windows clipboard via bracketed paste Mouse Paste (X11 style) Middle-click in pane Pastes Windows clipboard via bracketed paste Keyboard Paste Prefix + ] Pastes Windows clipboard into current pane Terminal Bypass Hold Shift while selecting or clicking Bypasses tmux mouse capture to use native Windows Terminal selection In the previous post on pane and window management, we looked at how to organize panes and navigate layouts quickly. Pairing those navigation keys with a sanitized clipboard pipeline eliminates the paper cuts of jumping between Windows editors and terminal shells inside WSL2.\n","permalink":"https://hanhpham.vercel.app/posts/tmux-wsl-clipboard-integration/","summary":"Piping tmux selections directly to clip.exe leaves trailing pane padding, broken prompt glyphs, and line ending mismatches in your clipboard. Here is a fast, sanitized copy-paste pipeline built for WSL2.","title":"WSL2 and tmux: Clean Clipboard Integration Without the Latency or Garbage Characters"},{"content":"Last week\u0026rsquo;s post took apart Huawei\u0026rsquo;s Ascend 950, a datacenter accelerator built under export controls. The natural reaction was \u0026ldquo;but how does it compare to the chip in my laptop?\u0026rdquo; So: Apple\u0026rsquo;s M4 family. Not because they compete, exactly. They don\u0026rsquo;t, and that\u0026rsquo;s the interesting part. One is the most sophisticated client SoC on the market; the other is a fleet component. Reading them against each other clarifies what \u0026ldquo;an AI chip\u0026rdquo; even means, because the two companies answered that question in almost perfectly opposite ways.\nOne scope note first: the M4 generation is a closed chapter. As of this week the Mac Studio ships with M5 Max and M5 Ultra (pre-orders open, delivery September 22), and there never was an M4 Ultra. Apple skipped straight past it. That fact will matter later.\nThe Core Takeaway: Autoregressive LLM decode is memory-bandwidth-bound, not compute-bound. An M4 Max ($2k laptop) hits 546 GB/s and runs an 8B model at ~90 tok/s. An Ascend 950DT (600W server card) hits 4.0 TB/s across an 8,192-chip fabric. They represent opposite answers to the same equation: local client privacy vs datacenter tokens-per-dollar.\nTwo theses The M4 and the Ascend 950 make opposite choices at every layer\nApple\u0026rsquo;s thesis: AI happens on the client, mostly in a fixed-function block, powered by a memory system shared with everything else. Inference runs where the data and the user already are, because it\u0026rsquo;s private, it\u0026rsquo;s free at the margin, and it works on a battery. Apple\u0026rsquo;s datacenter answer, Private Cloud Compute, is not a GPU server product; it\u0026rsquo;s custom Apple silicon behind a privacy architecture, and Apple refuses to say what\u0026rsquo;s in it.\nHuawei\u0026rsquo;s thesis: AI happens at fleet scale, in programmable engines, fed by dedicated memory, stitched into one fabric. Every design decision serves tokens-per-dollar across a datacenter, under sanctions.\nThe spec sheet, honestly M4 M4 Max M3 Ultra Ascend 950PR Ascend 950DT Product shape iPad / MacBook MacBook / Studio Mac Studio DC card DC card CPU 4P+6E 12P+4E 20-24P+8E Linx816 8C16T Linx816 8C16T Matrix engine 16-core ANE 16-core ANE 32-core ANE 32 Cube + 64 Vector 36 Cube + 72 Vector Peak matrix compute 38 TOPS (Apple) ~36 TFLOPS FP16 (est.) ~2x M4 Max (est.) 1,784 TFLOPS MXFP4 2,007 TFLOPS MXFP4 Formats FP16 (ANE), Metal GPU FP16 (ANE) FP16 (ANE) FP8/HiF8/MXFP8/MXFP4 same Memory 32 GB LPDDR5X 128 GB LPDDR5X up to 512 GB 128 GB HiBL 1.0 144 GB HiZQ 2.0 Bandwidth 120 GB/s 546 GB/s 819 GB/s 1.6 TB/s 4.0 TB/s Power battery battery / Studio ~250 W load (measured) ~600 W ~600 W Price ~$1k device from $1,999 (Studio) $8,099-11,699 ~$16k (reported) n/a Scaling beyond one none none none (M4 skipped Ultra) UB to 8,192 chips same Sources: Apple newsroom PRs (May 7, 2024; Oct 30, 2024; Mar 5, 2025) and archived spec pages; Ascend 950 white paper (2026); TrendForce for the 950PR price. Some rows need immediate unpacking:\n\u0026ldquo;38 TOPS\u0026rdquo; is Apple\u0026rsquo;s only published AI number, it belongs to the M4\u0026rsquo;s Neural Engine, and it\u0026rsquo;s an operations count at an unspecified precision (universally read as INT8). Apple publishes no FLOPS figures at all: not for the GPU, not for the ANE, not for any M-series chip, ever. The ~18 TFLOPS FP32 / ~36 TFLOPS FP16 figures floating around for the M4 Max\u0026rsquo;s 40-core GPU are third-party derivations (core count x clock x FLOPs-per-core, e.g. flopper.io), not measurements. For calibration, NVIDIA quotes the RTX 4090 at 165.2 TFLOPS of dense FP16 tensor math; an RTX 4090 draws about the wall power of four MacBook Airs. The Ascend\u0026rsquo;s 2,007 TFLOPS is MXFP4, a 4-bit tensor format with block scaling. Per-operation-per-watt, FP4 MACs are far cheaper than anything in the ANE. The 38 TOPS versus 2,007 TFLOPS gap is real, but it\u0026rsquo;s a gap between different questions: INT8 ops in a phone SoC\u0026rsquo;s power envelope versus FP4 tensor throughput in a 600 W card. The M3 Ultra row is there because the M4 generation has no ultra part. If you wanted a 512 GB Apple machine in 2025 and 2026, you bought the previous generation. That\u0026rsquo;s a thesis signal, not an accident. What Apple actually built: a secret in two and a half billion devices The Neural Engine has shipped in every Apple SoC since the A11 in 2017. It\u0026rsquo;s now in over 2.5 billion active devices. And until June this year, almost nobody outside Cupertino could even tell you what it is, because Apple documents nothing about it: no ISA, no driver interface, no programming guide. Core ML treats it as an opaque scheduling hint. There was no documented way to prove a computation ran on it.\nA June 2026 Georgia Tech paper changed that, by decompiling the private runtime, compiler, kernel driver, and firmware, then validating every claim against measurements on M1 and M5 silicon. What the ANE turns out to be:\nA fixed-function FP16 matrix accelerator. Each core is a 2D multiply tile backed by an 8-deep accumulator file, with an FP32-class accumulator per core. It is not a systolic array and not a small GPU; it\u0026rsquo;s a fixed-geometry array fed by per-tile DMA engines. Precision: FP16 compute, period. The datapath has no BF16 and no FP32 mode; those annotations are accepted by the compiler frontend and silently cast down. FP8 has an e4m3 lane in the element-type table that appears gated off. INT8 and INT4 exist as weight-streaming compression (int4 is a 16-entry palette lookup), not compute lanes. On the M4 generation, int8 weights mostly just save storage. ~2 MB of on-chip SRAM (M1) as the working set, interleaved across 64 banks. Past that, the engine tiles and streams from DRAM, which is exactly why decode-shaped workloads (streaming every weight for every token) cost it so much. Clocks rose from ~1.4 GHz (M1) to ~2.36 GHz (M5), and int8 compute runs at up to 2x the FP16 rate on the same array. Here\u0026rsquo;s the part that should reframe how you think about the M4: nobody runs LLMs on the ANE. MLX and llama.cpp, the two runtimes that define Apple-silicon inference, both drive the GPU via Metal; MLX\u0026rsquo;s supported devices are CPU and GPU, full stop. Apple\u0026rsquo;s own 2023 engineering note on deploying transformers to the ANE is a dead link, and the academic measurement work shows why: the ANE accepts a restricted op set through an opaque planner, placement is decided op-by-op in ways that can silently strand a whole model on the CPU, and recompilation per step costs seconds. Orion, bypassing Core ML through private APIs, got 170+ tok/s out of a tiny GPT-2 on an M4 Max, an impressive feat of plumbing that still proves the point: the ANE is where AI marketing lives, and the GPU is where AI workloads live.\nThe ANE is best understood as Apple\u0026rsquo;s bet on what inference would need in 2017, frozen into silicon: FP16 in, FP16 out, small models, fixed layers, maximum energy efficiency, zero developer surface. It\u0026rsquo;s a bet that paid off for Siri-scale workloads and was quietly routed around for everything else.\nContrast the DaVinci v3 Cube Core: programmable datapaths with native FP8, MXFP8, MXFP4, and a proprietary FP8 variant (HiF8), on-the-fly quantization at write-back, layout transforms in hardware, and a direct fusion path into vector cores for FlashAttention. Huawei is chasing the same energy prize Apple chases with fixed function, but through format flexibility instead of silicon inflexibility. The MXFP4-4x-BF16 claim in the white paper is the same claim NVIDIA makes for Blackwell. The ANE has no answer to FP4, and architecturally can\u0026rsquo;t: there\u0026rsquo;s no mantissa-flexible datapath to teach.\nThe numbers that actually matter: bandwidth and capacity The bandwidth gap is the whole story for decode workloads\nBoth companies agree on the load-bearing fact, and Apple\u0026rsquo;s own MLX team confirmed it with measurements: autoregressive decode is memory-bandwidth-bound. Tokens per second on a big model is roughly bandwidth divided by model bytes. This collapses the spec war into two numbers per chip:\nWorkload M4 Max (546 GB/s) M3 Ultra (819 GB/s) Ascend 950DT (4.0 TB/s) Qwen3-8B 4-bit decode ~90 tok/s (measured) ~135 tok/s (arith.) bandwidth not the limit 70B Q4 decode ~10 tok/s (arith.) ~15 tok/s (arith.) ~75 tok/s (arith.) 600B+ model, 8-bit doesn\u0026rsquo;t fit fits, ~30 tok/s (measured, 4-node cluster) one chip holds it; fabric pools more (M4 Max measured number from the apple-silicon-llm-bench harness, Qwen3-8B 4-bit, mean of 5 runs; 70B arithmetic is bandwidth / model-bytes and ignores overheads; the Ascend per-NPU serving numbers are Huawei\u0026rsquo;s own CloudMatrix-Infer paper on the previous-generation hardware: 1,943 decode tok/s per NPU on DeepSeek-R1.)\nThe capacity dimension cuts the same way. 128 GB is the M4 Max ceiling, and it\u0026rsquo;s gorgeous engineering, LPDDR5X at 546 GB/s inside a laptop. The Ascend 950DT puts 144 GB of HBM-class memory on one package and then does something the M4 generation structurally cannot: it lets 8,191 other packages read it.\nScaling: a fabric versus a Thunderbolt cable This is where the theses stop being mirror images and start being different species.\nApple\u0026rsquo;s scaling story, honestly told: there isn\u0026rsquo;t one, and Apple doesn\u0026rsquo;t sell one. Within a package, two dies fuse through UltraFusion at 819 GB/s (M3 Ultra) and 1.2 TB/s (M5 Ultra). Between machines, you have Thunderbolt 5: ~120 Gb/s raw, about 50-60 Gb/s real per Geerling\u0026rsquo;s measurements, no switches exist, so you cross-link machines point to point. His four-Studio cluster (1.5 TB of pooled memory, ~$40k) runs Kimi K2 Thinking, a trillion-parameter model, at ~30 tok/s using exo with RDMA over Thunderbolt 5. It\u0026rsquo;s a genuinely impressive engineering result, and it is the ceiling of Mac scale: four nodes, hobby budget, and the inter-node link is two orders of magnitude below the Ascend\u0026rsquo;s per-chip UB bandwidth.\nHuawei\u0026rsquo;s story: UB 2.0 gives every chip a 2 TB/s bidirectional pipe into a coherent Load/Store domain of up to 8,192 chips, with hardware collectives and IO dies that forward transit traffic without touching compute. It\u0026rsquo;s the exact architecture you\u0026rsquo;d design if your unit of competition were \u0026ldquo;tokens per day for a nation\u0026rsquo;s inference fleet\u0026rdquo;, which, given DeepSeek\u0026rsquo;s 160,000-chip order, it is.\nApple\u0026rsquo;s datacenter answer is different in kind: Private Cloud Compute, custom Apple silicon in hardened, stateless, publicly-auditable nodes, scaled by replication behind a load balancer rather than by coupling. Apple has never published a node spec, and the PCC design implies they consider tight coupling a liability (it is, for their threat model: stateless nodes can\u0026rsquo;t leak what they never held). Huawei scales out with a fabric; Apple scales out with copies; NVIDIA scales up and out with NVLink domains.\nPower and efficiency: where the thesis earns its keep The M4\u0026rsquo;s stack is absurdly efficient at small-scale inference, and the ANE is the reason even though the ANE usually isn\u0026rsquo;t the thing running the model. The Georgia Tech measurements put the ANE at a 1.8x energy advantage over the GPU on attention workloads, a ~0.9 W dispatch floor, and about 4-6 W under full compute load. Apple\u0026rsquo;s own on-device foundation model (3B parameters, quantized to 3.7 bits per weight) does 30 tok/s on an iPhone. MacBook Pro claims \u0026ldquo;up to 24 hours\u0026rdquo; of battery life, and LLM inference happens inside that envelope without a power cord in sight.\nThe Ascend 950PR is a 600 W card. Per-NPU decode in Huawei\u0026rsquo;s own paper is 1,943 tok/s on R1 at under 50 ms TPOT, which is excellent tokens-per-rack-unit, and their published training MFU on the previous generation (34.9%) trails NVIDIA\u0026rsquo;s typical 50-55%. Nobody buys an Ascend because it sips power; they buy it because it exists, it\u0026rsquo;s in stock (allocation-gated, but in stock), and it answers to no export license.\nThe honest symmetry: Apple wins tokens per watt and tokens per silence; Huawei wins tokens per dollar at fleet scale and tokens per rack under embargo. The M4 Max at ~65 W in a Studio (Apple publishes only the machine\u0026rsquo;s 480 W maximum continuous rating, not chip power) versus a 600 W Ascend card is not an Apple humiliation; it\u0026rsquo;s a different denominator.\nWhat each can\u0026rsquo;t do The M4 cannot scale. No fabric, no inter-node standard, no datacenter product you can buy, no FP8 or BF16 in its dedicated matrix engine, a 128 GB ceiling, and a matrix block that can\u0026rsquo;t be taught new formats. If your model outgrows one package, Apple\u0026rsquo;s answer is \u0026ldquo;buy another laptop\u0026rdquo;, or \u0026ldquo;wait for PCC\u0026rdquo;, or \u0026ldquo;that\u0026rsquo;s not our market\u0026rdquo;. The skipped M4 Ultra suggests even Apple sees the fused-die path as an M5-generation question.\nThe Ascend 950 cannot be personal. 600 W cards don\u0026rsquo;t run on battery, the software stack (CANN, ~80-90% operator coverage) is a decade behind CUDA\u0026rsquo;s gravity, the MFU gap is real, and the whole thing exists inside a sanctions wall that keeps getting rebuilt. It is also, conspicuously, not for sale to you.\nTwo theories of the AI chip It\u0026rsquo;s tempting to score this as a fight and declare NVIDIA the referee. That misses the shift. The market is bifurcating along exactly the line these two chips sit on: Benedict Evans frames token pricing as a market where some use cases work \u0026ldquo;just fine with a small, old, perhaps open source model that runs for \u0026lsquo;free\u0026rsquo; on your phone\u0026rdquo;, while others pay frontier datacenter rates. Apple built the chip for the first sentence. Huawei built the chip for the second. NVIDIA built a family for both and priced it accordingly.\nThe M4 is the most refined expression of the client thesis: a GPU that renders and infers, a fixed-function block for the predictable parts, memory shared with the OS, all of it invisible. The Ascend 950 is the rawest expression of the fleet thesis: formats, fabric, capacity, and volume production under embargo. One optimizes joules; the other optimizes sanctions.\nWhich one \u0026ldquo;wins\u0026rdquo; is the wrong question. The right one: for your workload, is the bottleneck energy and privacy (then it\u0026rsquo;s an M-series-shaped problem), or is it capacity and throughput at fleet scale (then it\u0026rsquo;s a fabric-shaped problem)? The chips stopped pretending to be interchangeable years ago. It\u0026rsquo;s the marketing that hasn\u0026rsquo;t caught up.\nSources Apple (primary): M4 PR (May 2024) and M4 Pro/Max PR (Oct 2024): core counts, bandwidth, 38 TOPS, N3E node. Mac Studio PR (Mar 2025): M3 Ultra, 819 GB/s, 600B-parameter claim. Private Cloud Compute (Jun 2024). Apple foundation models (Jun 2024). MLX on M5 (Nov 2025): decode is bandwidth-bound.\nANE reverse engineering: arXiv:2606.22283 (architecture), arXiv:2606.17090 (ANEForge), arXiv:2603.06728 (Orion, M4 Max), arXiv:2608.22110 (placement study).\nMeasurements: apple-silicon-llm-bench (M4 Max tok/s), Jeff Geerling\u0026rsquo;s 4-node cluster (Dec 2025), exo, MLX.\nHuawei side: all specs from the Ascend 950 white paper and the previous post; 950PR pricing via Spheron/TrendForce reporting; per-NPU serving numbers from CloudMatrix-Infer.\nEstimates are labeled in-line (GPU TFLOPS derivations, 70B tok/s arithmetic). Apple publishes no FLOPS, clocks, TDP, or die sizes for any M-series chip; if you see those stated as fact anywhere, they\u0026rsquo;re derived.\n","permalink":"https://hanhpham.vercel.app/posts/m4-vs-ascend-950/","summary":"One is a fixed-function FP16 block hidden inside a laptop chip Apple won\u0026rsquo;t document. The other is a programmable FP4 datacenter engine Huawei ships in 8,192-chip fabrics. Same goal, opposite theses: here\u0026rsquo;s what the M4 family actually is, and what the numbers do and don\u0026rsquo;t mean against the Ascend 950.","title":"Apple M4 vs Huawei Ascend 950: Two Theories of the AI Chip"},{"content":"tmux has ~200 default keybindings, and memorizing them is the wrong goal. The set that pays rent daily is maybe fifteen keys, plus three commands tmux deliberately left unbound that solve the problems the bound keys can\u0026rsquo;t. This post is that set, not the man page.\nThe 30-Second Cheatsheet: If you only memorize five bindings: prefix z (zoom pane), prefix ; (jump to last active pane), prefix % / prefix \u0026quot; (vertical/horizontal split), prefix c (new window), and prefix l (last window). For the three load-bearing commands tmux never bound (join-pane, move-pane, break-pane), jump to The four movers.\nThe mental model first, because everything below hangs off it: one tmux server hosts many sessions; each session has windows (the tabs in the status bar); each window has one or more panes (the rectangles on screen). Panes are cheap and disposable; windows are where work lives. Most beginners over-use panes and under-use windows; the reverse ages much better.\nOne housekeeping note: everything here is checked against upstream tmux 3.4 with an empty config (tmux -f /dev/null), and every command was tested live. If your ~/.tmux.conf (or a plugin) rebinds things, tmux list-keys is ground truth on your machine.\nSplits, and getting around once you\u0026rsquo;ve made them Key Action Notes prefix % Split side-by-side | shape: a vertical divider prefix \u0026quot; Split top-and-bottom Horizontal divider prefix ↑↓←→ Move between panes Repeatable; hold the arrow prefix o Next pane Cycles in layout order prefix ; Jump to the last active pane The edit→run→edit hop prefix q Flash pane numbers; press one to jump Fastest way across 4+ panes prefix z Zoom the current pane Same key restores. The focus key prefix { / prefix } Swap pane with previous / next Rearranges without moving yourself prefix x Kill pane (with confirmation) prefix \u0026amp; Kill window (with confirmation) The one trip-up everybody hits: % and \u0026quot; are named after the divider, not the result. % is a vertical bar → panes sit side by side. \u0026quot; is a double-quote → two lines stacked. Say it out loud once and it sticks.\nTwo of these deserve promotion into muscle memory ahead of everything else. prefix z is the single best quality-of-life key in tmux: one press gives the current pane the whole window (long build output, a big diff), one press gives your layout back. prefix ; is the heartbeat of the editor-plus-shell workflow: run tests in the bottom pane, watch them fail, press ; and you\u0026rsquo;re back on the line of code that caused it: no arrow-key counting, no q-lookup.\nThe four movers tmux never bound Default tmux can create panes and destroy panes, but it ships almost nothing for moving them around: join-pane, move-pane, and friends have no default bindings. They\u0026rsquo;re the difference between a tidy workspace and one where you close and re-split everything because a pane landed in the wrong window.\nThe trick that makes them pleasant is marking. prefix m toggles a mark on the current pane (the mark is global; it survives window switches), and prefix M clears it. Once a pane is marked, the move commands resolve to it automatically:\nCommand What it does join-pane Move the marked pane into the current pane\u0026rsquo;s window join-pane -h Same, forcing the new pane side-by-side move-pane -h -s {marked} Same idea, explicit: pull the marked pane next to the current one, even across windows break-pane (prefix !) Eject the current pane into its own new window move-window (prefix .) Re-index the current window: type a number, it moves swap-pane -U/-D ({ / }) Trade the current pane with its neighbors Gotcha worth knowing before it costs you ten minutes: in the swap-pane family (swap-pane, join-pane, move-pane), the marked pane is the implicit source. move-pane -t {marked} (which looks like a sensible \u0026ldquo;send my pane to the marked one\u0026rdquo; binding) silently does nothing once a mark exists, because source and target resolve to the same pane. The explicit -s {marked} form above is the one that works.\nThe daily rhythm this enables:\nEject and re-admit the long-running process. Your dev server doesn\u0026rsquo;t belong in the editor window, but you started it there. Press prefix ! and it\u0026rsquo;s now its own window with a status-bar tab you can check with prefix l (last-window). Want it back side-by-side later? Visit its window, prefix m to mark it, go home, prefix J\u0026hellip; except J isn\u0026rsquo;t bound by default. Which brings us to:\nA fifteen-line config that pays for itself # splits that inherit the current directory (defaults don\u0026#39;t) bind - split-window -v -c \u0026#34;#{pane_current_path}\u0026#34; bind \\ split-window -h -c \u0026#34;#{pane_current_path}\u0026#34; # pull the marked pane into the window you\u0026#39;re standing in bind J join-pane # ...side-by-side, when you care where it lands bind S move-pane -h -s {marked} # repeatable resizes without arrow-key origami bind -r H resize-pane -L 5 bind -r J resize-pane -D 5 bind -r K resize-pane -U 5 bind -r L resize-pane -R 5 set -g mouse on The -c \u0026quot;#{pane_current_path}\u0026quot; flag is small and outsized: every new pane opens in the directory you were already in, not ~. The resize bindings exist because the defaults are prefix C-arrow (one cell) and prefix M-arrow (five cells, both repeatable); fine, but H/J/K/L is faster to reach and the -r flag makes each press repeat until you move on.\nOne caution if you copy a popular vim-style pane-navigation config (binding h/j/k/l to direction keys): prefix l is last-window upstream, and those configs shadow it with \u0026ldquo;move right\u0026rdquo;. If you live in prefix l, rebind it explicitly (bind L last-window or similar); silently losing it is confusing for exactly as long as it takes you to forget the config change.\nLayouts: cycle, or jump straight to the one you want prefix Space cycles through tmux\u0026rsquo;s five built-in layouts: even-horizontal, even-vertical, main-horizontal, main-vertical, tiled. Cycling is fine for two panes; for more, jump directly:\nKey Layout Good for prefix M-1 even-horizontal Three panes in a row prefix M-2 even-vertical Stacked: editor / test / logs prefix M-3 main-horizontal Big pane on top, output below prefix M-4 main-vertical Editor left, stack right: the IDE look prefix M-5 tiled Grid: the dashboard prefix E spread current panes evenly Untangling a hand-resized mess (M- is Alt: prefix then Alt-4.) Don\u0026rsquo;t curate layouts by hand: resize when you need to (-r bindings above, or drag the border with the mouse), and let Space/M-4 do the rest. prefix C-o rotates panes through the layout if you want the same geometry with different contents.\nWindows and sessions: the part beginners skip Key Action prefix c New window prefix 0-9 Jump to window by number prefix p / prefix n Previous / next window prefix l Last window: ping-pong between two prefix , Rename window prefix . Move window to a new index prefix w Window tree, all sessions, searchable prefix s Session tree prefix ( / prefix ) Previous / next session prefix d Detach; everything keeps running Two habits turn this table from \u0026ldquo;aware\u0026rdquo; to \u0026ldquo;load-bearing\u0026rdquo;:\nName your windows, or lose the numbers. The status bar auto-renames windows to the running command, which means window 3 is called vim until you run make in it. prefix , and a two-letter name (ed, srv, log) makes prefix 3 land where you expect every time. Rename also stops the auto-rename, which is what you want.\nThe window is the unit of context, the pane is the unit of convenience. One window per task: editor panes here, the service stack there, the REPL in its own place. Inside a window, panes are free: split, zoom, throw them away. When a pane starts feeling important, that\u0026rsquo;s prefix ! asking to be a window.\nThe daily dozen If you keep only twelve keys, keep these:\nKey Job prefix % / prefix \u0026quot; Split prefix z Zoom in / out prefix ; Back to the last pane prefix q Pane numbers, jump prefix arrow Directional pane hop prefix Space / prefix M-4 Cycle / set layout prefix c, prefix n, prefix l Window: new, next, last prefix ! Pane → own window prefix m then prefix J Pull the marked pane here prefix d Detach Everything else is discoverable when you need it: prefix ? lists all bindings, prefix w shows every window in every session, and the three movers (join-pane, move-pane, break-pane) cover the rearranging that no default key does.\nThe natural next layer is copy mode: how to yank text out of a pane (and why vi-mode + a good copy-command matters more than any pane trick on this page). That\u0026rsquo;s a post of its own.\n","permalink":"https://hanhpham.vercel.app/posts/tmux-pane-window-management/","summary":"Most tmux guides dump the whole key table on you. These are the pane and window bindings that survive contact with a real workday: splits, navigation, layout, and the marked-pane tricks tmux never bound for you.","title":"tmux Panes \u0026 Windows: The Bindings That Actually Matter"},{"content":"On September 4, Bloomberg reported that DeepSeek plans to deploy at least 160,000 Huawei Ascend 950DT accelerators in a gigawatt-scale data center it is building in Inner Mongolia, one of the largest known Huawei-chip clusters anywhere. The chips are for inference only; DeepSeek still trains on NVIDIA hardware it managed to keep. Fulfilling the order could take more than a year, because Huawei\u0026rsquo;s 950DT output in 2026 is capped in the low hundreds of thousands by component shortages (memory, of all things).\nThe same season, Jensen Huang said NVIDIA\u0026rsquo;s China revenue has dropped to \u0026ldquo;essentially zero\u0026rdquo; and that NVIDIA has \u0026ldquo;largely conceded\u0026rdquo; the market to Huawei. Whether that\u0026rsquo;s a full retreat or a temporary posture is genuinely unclear (Washington approved H200 sales to Chinese firms in December 2025, and Beijing made sure almost none were delivered), but the direction is not: the market share NVIDIA doesn\u0026rsquo;t have anymore is being filled by a chip that, eighteen months ago, did not exist.\nThe Ascend 950 is interesting precisely because of what it is not. It is not a TSMC-class chip built on a bleeding-edge EUV node; it can\u0026rsquo;t be, because SMIC has no EUV machines and can\u0026rsquo;t buy one. It is a chiplet package on a DUV-only 7nm-class process, wrapped in self-developed HBM, glued by a homegrown interconnect into systems of up to 8,192 chips, and fed by a software stack that now claims day-zero support for the most popular open model on Earth. This post takes it apart: the process, the architecture, the package, the system, and an honest look at what\u0026rsquo;s still missing.\nThe Architecture at a Glance: The Ascend 950 is not an EUV miracle; it is an aggressive chiplet architecture on SMIC\u0026rsquo;s DUV 7nm-class node. Huawei compensated for lithography limits by engineering custom memory (HiZQ 2.0 HBM delivering 4.0 TB/s), a variable-width float format (HiF8), and an optical Unified Bus scaling to 8,192 chips in a single load/store domain.\nHow export controls became a design specification The timeline matters, because each US restriction converted directly into a Huawei design decision:\nDate Control Consequence for Huawei Oct 2022 A100/H100-class export ban Ascend 910 becomes stranded hardware Oct 2023 A800/H800 loopholes closed No compliant NVIDIA parts above ~half H100 2024 NVIDIA ships H20, built to comply The \u0026ldquo;legal\u0026rdquo; chip becomes China\u0026rsquo;s default Dec 2024 HBM banned (ECCN 3A090.c: any stack \u0026gt;2 GB/s/mm² of bandwidth density) Every production HBM stack qualifies: Huawei must build its own Sep–Nov 2024 TSMC dies found in 910B/910C teardowns; TSMC halts supply (~$500M penalty reported) Stockpiled TSMC 7nm dies (~2.9M, per SemiAnalysis) become a dwindling asset Apr 2025 H20 license requirement ~$4.5B NVIDIA write-off; China\u0026rsquo;s default chip gone Aug 2025 15% China-revenue deal; H20 briefly returns Sep 2025: Beijing\u0026rsquo;s CAC discourages NVIDIA purchases Nov 2025 B30A (Blackwell-class for China) blocked The ceiling on legal NVIDIA silicon freezes Dec 2025 – May 2026 H200 approved; Beijing stalls deliveries, then formally bars big buyers Huang: China revenue \u0026ldquo;essentially zero\u0026rdquo; Sources: BIS press releases, CSIS analysis, Tom\u0026rsquo;s Hardware\u0026rsquo;s Huawei teardown coverage, SemiAnalysis, Reuters, FT via China AI Dispatch.\nThe December 2024 HBM ban is the one most people undersell. HBM is where bandwidth lives, bandwidth is what inference runs on, and until that day every Chinese accelerator used Samsung or SK Hynix stacks. Cutting HBM off didn\u0026rsquo;t slow Huawei\u0026rsquo;s roadmap: it added a product line to it. The HiBL 1.0 and HiZQ 2.0 memory stacks on the Ascend 950 exist because the alternative was no memory at all.\nBy 2025 the strategy had stopped being \u0026ldquo;evade the controls\u0026rdquo; and become \u0026ldquo;build a parallel stack\u0026rdquo;: SMIC for silicon, HiSilicon for HBM, the Unified Bus for interconnect, CANN for software. The Ascend 950 is the first chip designed end-to-end inside that constraint set. That\u0026rsquo;s the lens for everything below.\nSMIC N+3: a 7nm-class node without EUV, measured at last The Ascend 950\u0026rsquo;s compute dies are widely reported to be made on SMIC\u0026rsquo;s N+3 process (analyst reporting: TrendForce, citing EE Times China and Eastmoney; Huawei itself only says \u0026ldquo;fully self-controlled manufacturing\u0026rdquo;). Whatever the fab assignment, N+3 is real and well-characterized now, because it ships in Huawei\u0026rsquo;s Kirin 9030 smartphone SoC, and two teardown firms have cut it open.\nTechInsights confirmed in December 2025 that the Kirin 9030 is built on N+3, a \u0026ldquo;scaled extension\u0026rdquo; of SMIC\u0026rsquo;s 7nm N+2, and explicitly not a 5nm-class node. Then SemiAnalysis\u0026rsquo;s STEEL teardown in June 2026 put numbers on it, and the numbers are more interesting than the marketing:\nMetric SMIC N+3 TSMC N6 Intel 18A (shipping) Min metal pitch (M0) 32.5 nm ~40 nm 36 nm Transistor density 113.4 MTr/mm² 107.7 MTr/mm² ~38% higher (normalized, HD library) Patterning of critical layers SAQP (M0, fin), SADP (M1/M2) EUV EUV Transistor FinFET, 2 fins (depominated) FinFET GAA RibbonFET Backside power No No Yes (PowerVia) Lithography DUV immersion only EUV + DUV EUV + DUV Cell height is 228 nm (5.7-track), with contact-over-active-gate and single diffusion breaks: aggressive DTCO, in other words, squeezing density out of design rules rather than lithography. The 32.5 nm M0 pitch is tighter than what Intel ships on 18A, and density edges past TSMC N6, a node that uses EUV.\nRead those comparisons carefully, though. The pitch headline is one SemiAnalysis itself calls cherry-picked: Intel supports 32 nm on 18A and ships looser high-performance libraries. And density numbers only compare within one methodology: N+3\u0026rsquo;s 113.4 MTr/mm² and the often-quoted \u0026ldquo;18A = 238 MTr/mm²\u0026rdquo; are on different bases. The honest summary: N+3 reaches TSMC N6-class density the hard way: more masks, more overlay sensitivity, worse cost and yield; and it\u0026rsquo;s nowhere near N5/N4, let alone 18A, on efficiency. TechInsights flags that BEOL yield on SAQP layers can fall off a cliff once overlay budgets are exceeded.\nSemiAnalysis projects N+4 landing near TSMC N5-class density (~138 MTr/mm²) and N+5 near 18A-class (~164, with backside contacts); projections, not measurements. For the Ascend 950, the practical takeaway is different: the chip doesn\u0026rsquo;t need leading density, because it compensates at the system level. What it needs is adequate density, available at volume, without anyone\u0026rsquo;s export license. That\u0026rsquo;s what N+3 is.\nDaVinci v3: two specialized cores instead of one general one The Ascend 950 runs the third generation of Huawei\u0026rsquo;s DaVinci architecture, documented in Huawei\u0026rsquo;s own Ascend 950 NPU Architecture White Paper (2026), the primary source for everything in this section.\nDaVinci\u0026rsquo;s founding bet (per the Hot Chips 2019 and HPCA 2021 papers) was that AI compute is mostly matrix math plus a tail of element-wise work, so you build a Cube Core for the matrices and a Vector Core for everything else, rather than one SIMT machine doing both badly. v3 keeps that separation and sharpens it: the full chip has 36 AI subsystems, each one 1 Cube Core + 2 Vector Cores.\nOne AI subsystem of DaVinci v3: the tile repeats 36 times per 950DT package\nThe Cube Core: where the FLOPS live, and where the bits shrink The Cube Core is the tensor engine, and v3\u0026rsquo;s headline is a new precision ladder:\nFormat 950DT peak (full config) Relative TF32 273 TFLOPS 0.5× BF16 / FP16 547 TFLOPS 1× FP8 / MXFP8 / HiF8 1,034 TFLOPS 2× INT8 1,034 TOPS 2× MXFP4 2,007 TFLOPS 4× Two details matter more than the peak numbers:\nOn-the-fly quantization at L0C write-back. The 256 KB L0C accumulator buffer doesn\u0026rsquo;t just hold FP32 accumulation results: it converts them on the way out to BF16/FP16/FP8/INT8, and re-lays them out (NZ→ND) in the same operation. If your next consumer wants FP8, the 32-bit intermediate never pays DRAM bandwidth for being 32-bit. This is exactly the \u0026ldquo;fused online quantization in hardware\u0026rdquo; that DeepSeek-V3\u0026rsquo;s training report explicitly asked chip designers for (§3.5) after training a 671B MoE in FP8. Huawei read the same literature.\nBigger L0C for FlashAttention. A 256 KB accumulator supports more aggressive tiling of attention kernels, and Huawei claims 1.5–2× single-core FlashAttention performance over the previous generation via the fused Cube-Vector path (next section).\nThe Vector Core: fixed the bottleneck the last generation created Previous Ascends had a dirty secret: the Vector Cores were weak enough that non-matrix work (Softmax, GELU, normalization: FlashAttention\u0026rsquo;s back half) starved the Cube Cores. v3 fixes it with a redesign:\nRegister-based, dual-issue, out-of-order SIMD: a real vector machine now, not a vector coprocessor, with a RegFile between the Unified Buffer and the ALUs FP16/FP32 per-core throughput up 100%, plus native BF16 and a full set of conversion instructions Microcoded hot functions: Softmax and GELU get dedicated datapath treatment to keep the tensor ALUs fed SIMD + SIMT on the same core The genuinely novel bit is the programming model Huawei calls \u0026ldquo;new homogeneous\u0026rdquo; SIMD/SIMT hybrid. Each unit of work (a Vector Function) can run in either mode:\nSIMD mode (the default): regular element-wise work, dual-issue, high throughput SIMT mode (the escape hatch): irregular access patterns (gather/scatter, hash inserts, branching) get per-thread addressing like a GPU, without giving up the SIMD path for everything else NVIDIA\u0026rsquo;s bet is SIMT with tensor cores bolted on; DaVinci\u0026rsquo;s is SIMD-first with a SIMT lane for the awkward 10%. Neither is wrong; they optimize for different distributions of work. For recommendation systems and multimodal preprocessing (real Huawei workloads), the SIMD-first split is defensible.\nNDDMA and painless plumbing Two quieter changes cut the cost of writing kernels: NDDMA, a DMA engine that performs up to 5-dimensional layout transforms (NCHW↔NHWC-style rearrangements, transposes) in flight while coalescing small reads into 128-byte sectors; and a BufferID synchronization mechanism that replaces Ascend\u0026rsquo;s famously fiddly set_flag/wait_flag discipline with something that looks like a mutex (get_buf/rel_buf). Neither shows up on a datasheet. Both show up in developer productivity, which is where Ascend\u0026rsquo;s real deficit has always been.\nHiF8: an 8-bit float with (almost) FP16\u0026rsquo;s range The most technically interesting part of the launch is a data format. Huawei published an academic version (\u0026ldquo;Ascend HiFloat8 Format for Deep Learning\u0026rdquo;, 2024), and the white paper confirms it\u0026rsquo;s native in the 950\u0026rsquo;s Cube Cores.\nHiF8 trades mantissa precision for range as magnitude shrinks: a tapered, cone-shaped encoding\nThe problem with FP8 is that E4M3 spends its 8 bits badly for neural-network data: 18 powers of two of dynamic range, and out-of-range activations either saturate or need rescuing. The Microscaling formats (MX, now an industry standard) fix this by attaching a shared 8-bit scale factor to each 32-element block: that\u0026rsquo;s what MXFP8 and MXFP4 are. HiF8 makes a different trade: a variable-width exponent. A prefix code (\u0026ldquo;Dot\u0026rdquo;) declares how wide the exponent field is; the exponent is stored sign-magnitude with one hidden bit so the ranges of different widths never overlap; the mantissa gets whatever\u0026rsquo;s left. Precision is highest near |x| ≈ 1 and tapers off: 7 binades with 3 mantissa bits, 8 with 2, 16 with 1, per the paper. The result: 38 combined exponents (2⁻²² to 2¹⁵) versus E4M3\u0026rsquo;s 18, approaching FP16\u0026rsquo;s 40, in 8 bits, with no extra scale factor. Four special values (zero, NaN, ±Inf), no negative zero.\nWhy bother, when MX exists? Two reasons. First, the MX scale factor is real overhead: an extra 8 bits per 32 elements (25% overhead on FP4 storage), a second operand in every GEMM, and a whole microarchitecture of scale-handling. HiF8 needs none of it. Second, range: Huawei\u0026rsquo;s claim, supported by the paper\u0026rsquo;s simulations, is that HiF8\u0026rsquo;s dynamic range is wide enough to keep both forward and backward passes in 8-bit tensors, where E4M3-based recipes need 16-bit elsewhere in the step. That\u0026rsquo;s the same insight behind Intel\u0026rsquo;s trillion-token FP8 work (arXiv:2409.12517): the failures at scale are range failures. If the claim holds in production training, HiF8 is a genuinely elegant piece of engineering, and it says something that the format exists because one company controls the whole stack, format to compiler to silicon.\nThe package: two AI dies, two IO dies, and memory that doesn\u0026rsquo;t exist anywhere else The Ascend 950 is not one die. It\u0026rsquo;s a chiplet assembly: 2 AI Dies + 2 IO Dies + 8 (950PR) or 4 (950DT) on-package memory stacks, joined by Huawei\u0026rsquo;s Clink die-to-die links and memory interfaces into a single UMA: one address space, hardware-coherent L2 across both AI dies, software none the wiser.\nAscend 950 package topology: the IO dies hold every SerDes, so scale-up transit traffic never touches the compute dies\nThe two-SKU split is the package used as a product strategy:\nAscend 950PR (Q1 2026) Ascend 950DT (Q4 2026) Role Prefill \u0026amp; Recommendation Decode \u0026amp; Training AI subsystems (Cube + 2×Vector) 32 (bin: 28) 36 (bin: 32, 28) MXFP4 / FP8 peak 1,784 / 919 TFLOPS 2,007 / 1,034 TFLOPS On-package memory 128 GB HiBL 1.0 @ 1.6 TB/s (bin: 112 GB @ 1.4) 144 GB HiZQ 2.0 @ 4.0 TB/s (bin: 96 GB) L2 cache 128 MB (bin: 112) 128 MB AI CPU Linx816, up to 8C16T Linx816, up to 8C16T UB interconnect 2.0 TB/s bidirectional 2.0 TB/s bidirectional (All full-config numbers from the white paper; the bins are the same silicon with redundant blocks disabled, since binning isn\u0026rsquo;t optional on a SAQP process.)\nWhy chiplets? The same reason everyone else does them: yield, amplified by the process. On N+3\u0026rsquo;s multi-patterned metal layers, a big monolithic die would be a yield catastrophe. Small AI dies plus relaxed-node IO dies plus binned SKUs is how you get a large chip out of a difficult node. NVIDIA does chiplets for HBM capacity and modularity; Huawei does them for survival, and the survival version looks structurally similar.\nThe IO dies are the quietly clever part. All 72 HiLink SerDes lanes (grouped as 18 ×4 ports at up to 112 Gbps), the PCIe 5.0 controller, and the two 400G Ethernet ports hang off the IO dies, and so does the UB On-Chip Switch. Transit traffic flowing through this chip to elsewhere in the rack gets forwarded on the IO die without touching a compute die or consuming DRAM bandwidth. It\u0026rsquo;s a switch built into the endpoint, and it\u0026rsquo;s the architectural foundation for the scaling story below.\nAlso worth naming: the memory. HiBL 1.0 (950PR) is Huawei\u0026rsquo;s self-developed, cost-optimized HBM: the keynote positioned it as cheaper than HBM3E; the shipped 950PR cards (Atlas 350, unveiled March 20, 2026) carry 112 GB of it, with teardown reporting of no Samsung, SK Hynix, or Micron content. HiZQ 2.0 (950DT) is the HBM4-class tier: 4 TB/s per chip. Reporting (Convequity, SemiAnalysis-adjacent) suggests the DRAM dies come from Swaysure with Huawei\u0026rsquo;s own base die (plausible, unconfirmed at fab level). Either way: the December 2024 HBM ban got answered in fourteen months, at volume, inside a shipping product.\nThe system bet: Unified Bus and the 8,192-chip SuperPoD A chip two generations behind NVIDIA\u0026rsquo;s best cannot win alone. So Huawei changed the unit of competition (from chip to system), and Unified Bus (UB) 2.0 is the instrument.\nOne UB fabric at three scales: the same memory semantics from package to SuperPoD\nUB is a full interconnect stack, and its spec was publicly released in September 2025 with an open-source pledge attached:\nUB Memory: synchronous Load/Store/Atomic across a shared address space of up to 128 TB. Remote NPU memory behaves like local memory; there is no \u0026ldquo;copy then compute\u0026rdquo; step. URMA: asynchronous one-sided copies over queue pairs (\u0026ldquo;Jetties\u0026rdquo;), with two transport layers: RTP (end-to-end reliable retransmission, 4 ports) and CTP (lightweight, 9 ports). CCU: a Collective Communication Unit that hardware-executes Broadcast, ReduceScatter, AllGather, AllReduce and All2All off the compute cores, using its own memory slice and reduce units. Collectives stop stealing AI cores and NoC bandwidth. UBoE: UB tunneled over standard Ethernet, so a Huawei cluster can hop onto any existing datacenter fabric without new switches. Latency: 2.1 μs across the fabric, per the keynote; not NVLink-class, but the claim is that memory semantics (no copy, no handoff) beat raw latency at system scale. The scale targets are where this stops being incremental. The previous generation\u0026rsquo;s CloudMatrix 384 already networked 384 chips; the Atlas 950 SuperPoD networks 8,192 of them: 160 cabinets, ~1,000 m², all-optical, 1,152 TB of on-package memory, 16.3 PB/s of aggregate interconnect, 8 EFLOPS FP8 / 16 EFLOPS FP4, targeted for Q4 2026. Huawei\u0026rsquo;s own comparison: 62× the interconnect bandwidth and 15× the memory of an NVIDIA NVL144 domain shipping the same period. Scale it again and the Atlas 950 SuperCluster crosses 520,000 chips at a zettaflop of FP4.\nIs \u0026ldquo;weaker chip, vastly bigger fabric\u0026rdquo; a legitimate strategy? It\u0026rsquo;s arithmetic. If your chip is ~2× behind in FP8 and ~2.5× behind in memory bandwidth per chip, but you can network 8,192 of them with coherent memory semantics while your competitor\u0026rsquo;s scale-up domain is 72–144 GPUs, the system-level gap closes and partly inverts, at the price of floor space, optics, power, and software complexity. NVIDIA knows this; NVL144 and its roadmap are the same logic with better per-chip economics. The difference is that Huawei must win at this game, and so built UB (72 lanes per chip, port-multiplexed across scale-up, PCIe, and Ethernet) as a first-class citizen rather than a proprietary garnish.\nMemory: the whole ladder, and the wall it\u0026rsquo;s aimed at Put the whole hierarchy side by side and the design intent is legible:\nSix levels from ALU to 128 TB: note the on-package memory tier, which is where the HBM ban bit hardest\nChip Memory Bandwidth Per-chip FP8 TDP NVIDIA H20 (2024, legal China chip) 96 GB HBM3 4.0 TB/s 148 TFLOPS 400 W Ascend 950PR 112–128 GB HiBL 1.0 1.4–1.6 TB/s ~0.8–0.9 PFLOPS ~600 W Ascend 950DT 144 GB HiZQ 2.0 4.0 TB/s ~1.0 PFLOPS n/a NVIDIA B200 (2025) 192 GB HBM3e 8.0 TB/s 4.5 PFLOPS (dense) 1,000 W The L2 cache deserves its own note, because it\u0026rsquo;s where Huawei spent transistor budget that a TSMC node would have spent elsewhere: 128 MB, chiplet-wide, coherent across dies, 512-byte lines split into four 128-byte sectors, per-way locking, and programmable cache hints (a producer task can mark its output non-allocate so streaming junk never evicts the KV-cache-adjacent data the next task needs). Huawei claims 2× the previous generation on random and small-packet access patterns at equal bandwidth; for inference serving, where L2 hits on KV data are money, that\u0026rsquo;s a real number.\nPer-chip against NVIDIA, the gap is what it is: roughly 4–4.5× behind B200 in FP8, ~2× behind in bandwidth, ~1.5× behind in capacity per chip. But note what the ladder does deliver: 144 GB and 4 TB/s is far past the H20 that served China for two years, and it exists at all only because Huawei was forced to build memory. The wall the ladder is aimed at is the KV cache, and there, capacity per system (1,152 TB per SuperPoD, plus pooled CPU memory over UB) is the metric that matters, not per-chip bandwidth.\nWhat it means for how inference gets built: PD separation as silicon The 950\u0026rsquo;s two SKUs (Prefill \u0026amp; Recommendation, Decode \u0026amp; Training) are a hardware encoding of an idea the systems community has been converging on for three years: prefill and decode are different workloads and should run on different machines. Prefill (processing the prompt) is compute-bound; decode (generating tokens) is memory-bandwidth-bound. Splitwise (ISCA \u0026lsquo;24) showed 1.4× throughput at 20% lower cost from splitting them; DistServe (OSDI \u0026lsquo;24) formalized the goodput math; Mooncake (FAST \u0026lsquo;25) runs it in production for Kimi with a pooled KV cache.\nHuawei\u0026rsquo;s CloudMatrix-Infer paper (June 2025) is the missing production-scale proof: a 384-NPU UB fabric where prefill, decode, and KV-cache pools scale independently, with expert parallelism up to EP320. On DeepSeek-R1 it sustains 6,688 prefill tokens/s and 1,943 decode tokens/s per NPU (modest per-chip, multiplied by 384). The 950 series turns that architecture into product: buy PR cards for your prefill fleet and DT cards for your decode fleet, wired into one UB memory domain, with STARS 2.0 scheduling compute and CCU handling the collectives.\nThis is also the honest frame for the DeepSeek order. 160,000 950DTs (the decode-and-training SKU) for a datacenter that will serve inference for a 1.6T-parameter open-weights model whose CANN port landed on day zero. Reporting (SemiAnalysis, via InfoQ) claims V4 and the 950DT were co-designed: the model\u0026rsquo;s decode stage shaped around HiZQ\u0026rsquo;s 4 TB/s. Whether or not that\u0026rsquo;s exactly right, the direction is real: the hardware-software co-evolution that made CUDA\u0026rsquo;s moat is being attempted, at enormous scale, from the other side.\nThe honest scorecard Where Huawei is genuinely strong:\nSystem architecture. UB\u0026rsquo;s memory semantics, the IO-die transit switch, hardware collectives, and 8,192-chip coherent domains are a coherent answer to the \u0026ldquo;per-chip deficit\u0026rdquo; problem, arguably a year ahead of NVIDIA\u0026rsquo;s scale-up ambition in topology, if not in per-node efficiency. Vertical integration under duress. Process (SMIC N+3), memory (HiBL/HiZQ), interconnect (UB), compiler (CANN, operator coverage reported up from ~30% to 80–90% since 2024), and now format design (HiF8). No other sanctioned company has pulled this off. Product-market fit inside China. The 950PR shipped on schedule (Q1 2026); Huawei targets ~750,000 950-series units this year; ByteDance alone has reportedly booked half the production line. IDC counted 812,000 Ascend cards shipped in China in 2025. The demand exists and is prepaid. Where the gaps are real:\nPer-chip efficiency. ~4× behind B200 in FP8 per chip, and power tells the same story: the shipped 950PR does ~1.56 PFLOPS FP4 at 600 W where B200 does 4.5 PFLOPS FP4 (dense) at 1,000 W. N+3 costs you performance per watt, and no amount of chiplet trickery hides that. MFU and software gravity. The best published training run on Ascend hardware (a 1.6T-parameter post-training on ~1,000 910Cs) achieved 34.9% MFU against 50–55% typical on NVIDIA. CANN has closed most of the operator gap but not the ecosystem gap: CUDA is a decade of Stack Overflow answers, and every PyTorch workflow defaults to it. Yield and supply. SAQP yields are the industry\u0026rsquo;s known unknown; 2026 production (\u0026ldquo;low hundreds of thousands\u0026rdquo; of 950DT-capable output against a 750,000-unit target and a 4.2M-chip domestic demand estimate) means allocation, not purchase, is the bottleneck. DeepSeek needed Beijing\u0026rsquo;s help to jump the queue. Training, still. The 160,000-chip DeepSeek order is for inference. The training fleet remains NVIDIA (H800s) plus whatever Ascends can be spared. The 950DT is designed for training; nobody has demonstrated frontier-class pretraining on it yet. That proof point is the whole game, and it hasn\u0026rsquo;t been played. The fair summary is not \u0026ldquo;Huawei caught NVIDIA.\u0026rdquo; It\u0026rsquo;s that Huawei removed NVIDIA\u0026rsquo;s leverage: for the workloads that dominate Chinese AI spend today (serving large open models at extreme scale), the domestic stack is now sufficient, cheaper per token, and immune to export policy. Compelling hardware everywhere was never the requirement.\nWhat it means for the rest of the world Three things, in ascending order of importance.\nFirst, export controls changed the optimization problem instead of ending it. Cut off from EUV, China\u0026rsquo;s answer was DUV multi-patterning plus chiplets plus system-scale compensation: a worse chip and, for large-model serving, a competitive system. Cut off from HBM, Huawei built memory. Every restriction so far has been answered inside roughly eighteen months, at shipping volume. The controls still hurt (the power and yield gaps are permanent tax), but \u0026ldquo;deny\u0026rdquo; and \u0026ldquo;delay\u0026rdquo; are turning out to be different words.\nSecond, the unit of competition has moved from chip to system, and the interconnect won. NVIDIA\u0026rsquo;s NVLink, Google\u0026rsquo;s ICI, and now Huawei\u0026rsquo;s UB are the same admission: memory-bandwidth economics at cluster scale beat single-chip FLOPS. Watch UB specifically: it has published specs, an open-source pledge, and Ethernet interoperability (UBoE). Interconnect standards are how regional platforms go global; PCIe got there, InfiniBand got there, and a UB that works over any Ethernet switch is not a crazy candidate.\nThird, the low-precision end is now a three-way format race. NVIDIA ships MXFP4/FP8, Huawei ships MXFP4/FP8 plus HiF8, and the research literature (MX, FP4 training, trillion-token FP8) has made sub-8-bit training respectable. Whoever\u0026rsquo;s tensor cores most natively express what the models actually need (block scaling, wide dynamic range, in-flight quantization) wins efficiency per dollar for the next generation. The Ascend 950\u0026rsquo;s format roster is a shot across that bow from a company nobody had on their format-race bingo card.\nThe scoreboard to watch for the rest of 2026: does the 950DT actually ship at volume in Q4? Does anyone publish frontier-class pretraining results on it? Does CANN\u0026rsquo;s open-source deadline hold? And when the Ascend 960 lands next year with 288 GB and 9.6 TB/s per chip, the question stops being \u0026ldquo;can China build an AI chip?\u0026rdquo; and becomes \u0026ldquo;how much of the world\u0026rsquo;s inference can it serve?\u0026rdquo;, which is a very different question, with a much larger answer.\nSources Primary (Huawei):\nAscend 950 NPU Architecture White Paper (2026): chiplet topology, DaVinci v3, HiF8, UB 2.0, STARS 2.0, CCU, memory hierarchy Eric Xu keynote, Huawei Connect (Sep 18, 2025): roadmap, HiBL/HiZQ, Atlas 950/960 SuperPoD and SuperCluster figures Ascend HiFloat8 Format for Deep Learning (arXiv:2409.16626) Serving LLMs on Huawei CloudMatrix384 (arXiv:2506.12708) Teardowns and analysis:\nSemiAnalysis, \u0026ldquo;Is SMIC N+3\u0026rsquo;s Metal Pitch Smaller than Intel 18A\u0026rsquo;s?\u0026rdquo; (Jun 2026): all N+3 quantitative figures TechInsights, \u0026ldquo;SMIC N+3 Confirmed: Kirin 9030\u0026rdquo; (Dec 2025) and process-flow analysis (Aug 2026) SemiAnalysis, \u0026ldquo;Huawei AI CloudMatrix 384\u0026rdquo; (Oct 2025) Tom\u0026rsquo;s Hardware, Ascend roadmap and UB deep-dives (2025) News:\nBloomberg, \u0026ldquo;DeepSeek Plans Big Huawei AI Chip Order\u0026rdquo; (Sep 4, 2026) Reuters: DeepSeek V4 on Huawei chips (Apr 2026), US clears H200 for ten Chinese firms (May 2026) TrendForce: Atlas 350 debut (Mar 2026), 950DT pull-forward (Jun 2026) Academic context:\nFP8 formats: arXiv:2209.05433 · Microscaling: arXiv:2310.10537 · DeepSeek-V3: arXiv:2412.19437 · FP4 training: arXiv:2502.20586 · Trillion-token FP8: arXiv:2409.12517 PD separation: Splitwise, DistServe, Mooncake · Roofline survey: arXiv:2402.16363 Corrections welcome: figures marked as analyst claims are attributed, and the white paper remains the authority for everything on-chip.\n","permalink":"https://hanhpham.vercel.app/posts/huawei-ascend-950-davinci-v3/","summary":"DeepSeek just ordered 160,000 Huawei Ascend 950DT chips: a chip built on a DUV-only process, with self-developed HBM and a homegrown interconnect, because export controls left no other choice. Here\u0026rsquo;s how the thing actually works, and where it still falls short.","title":"Huawei's Ascend 950 \u0026 DaVinci v3: How China Built a Competitive AI Chip Without EUV"},{"content":"This is a separate problem from wildcard joins and OR chains, but the same neighborhood: you\u0026rsquo;re running a script with several steps, maybe calling into other stored procedures, and one of them might fail partway through. What actually happens to your data depends entirely on how you\u0026rsquo;ve wired up error handling, and the default behavior is not what most people assume.\nI tested three variants against the same failure: a nested procedure that inserts a row, hits a divide-by-zero error, then (in source code, never in practice, because the error stops it) tries to insert a second row. Here\u0026rsquo;s the real ledger output from each run: not paraphrased, not estimated.\nThe 10-Second Takeaway: In SQL Server, BEGIN TRAN without SET XACT_ABORT ON does not roll back on runtime error by default; it silently commits partial data. For trustworthy multi-step migrations, always prepend SET XACT_ABORT, NOCOUNT ON; and use a structured BEGIN TRY ... BEGIN TRAN ... COMMIT TRAN END TRY BEGIN CATCH IF @@TRANCOUNT \u0026gt; 0 ROLLBACK TRAN; THROW; END CATCH block.\nScenario 1: plain BEGIN TRAN, no error handling SET XACT_ABORT OFF; BEGIN TRAN; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: before EXEC\u0026#39;); EXEC #ChildStep; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: after EXEC\u0026#39;); COMMIT; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: COMMIT succeeded\u0026#39;); Ledger:\nouter: before EXEC child: before error child: after error outer: after EXEC outer: COMMIT succeeded Every statement ran anyway, including the one after the error inside the child procedure, and the final COMMIT. Nothing crashed. Nothing rolled back. It just quietly kept going. This is SQL Server\u0026rsquo;s default: most runtime errors are not fatal to the batch, let alone the transaction.\nScenario 2: TRY/CATCH added SET XACT_ABORT OFF; BEGIN TRY BEGIN TRAN; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: before EXEC\u0026#39;); EXEC #ChildStep; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: after EXEC\u0026#39;); COMMIT; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: COMMIT succeeded\u0026#39;); END TRY BEGIN CATCH INSERT INTO #Ledger (Step) VALUES (\u0026#39;CATCH fired, XACT_STATE=\u0026#39; + CAST(XACT_STATE() AS varchar(3))); IF XACT_STATE() \u0026lt;\u0026gt; 0 ROLLBACK; INSERT INTO #Ledger (Step) VALUES (\u0026#39;CATCH: after rollback check\u0026#39;); END CATCH; Ledger:\nCATCH: after rollback check CATCH fired immediately this time, better. But look at what\u0026rsquo;s missing: the diagnostic line logged before the ROLLBACK (\u0026ldquo;CATCH fired, XACT_STATE=\u0026hellip;\u0026rdquo;) got wiped out by that same rollback, because it was inserted into #Ledger inside the same transaction that then got undone. Only the line logged after rolling back survived.\nScenario 3: XACT_ABORT ON, rollback first SET XACT_ABORT ON; BEGIN TRY BEGIN TRAN; ... COMMIT; END TRY BEGIN CATCH IF XACT_STATE() \u0026lt;\u0026gt; 0 ROLLBACK; -- always first SELECT ERROR_NUMBER(), ERROR_PROCEDURE(), ERROR_LINE(), ERROR_MESSAGE(); -- diagnostics after END CATCH; Ledger:\nCATCH: state=-1 proc=#ChildStep line=5 msg=Divide by zero error encountered. Everything before the error was cleanly undone. Only the diagnostic log written after rolling back survived, and it captured the exact failing line and procedure.\nGetting here needed one fix mid-testing: my first attempt logged the error before rolling back, same as scenario 2\u0026rsquo;s mistake, and that log line vanished too. Once a transaction is doomed (hit an error severe enough that it can no longer be committed), nothing else it does can be written until it\u0026rsquo;s rolled back. Rollback has to come first in the CATCH block, always.\nXACT_ABORT ON is the setting doing the real work here. It\u0026rsquo;s what turns \u0026ldquo;an error happened somewhere in this chain\u0026rdquo; into \u0026ldquo;the whole transaction is instantly, automatically undone\u0026rdquo;, with no dependency on execution ever reaching your ROLLBACK line. Without it, as scenario 1 showed, a runtime error can be entirely survivable from the engine\u0026rsquo;s point of view, even though it\u0026rsquo;s exactly the kind of thing you\u0026rsquo;d want to stop everything for.\nThe pattern behind scenario 3, in full, is the one worth keeping around:\nSET XACT_ABORT ON; BEGIN TRY BEGIN TRAN; ... -- your steps here COMMIT; END TRY BEGIN CATCH IF XACT_STATE() \u0026lt;\u0026gt; 0 ROLLBACK; -- always first SELECT ERROR_NUMBER(), ERROR_PROCEDURE(), ERROR_LINE(), ERROR_MESSAGE(); -- diagnostics after END CATCH; Try it yourself Self-contained: a local temp table for the ledger and a local temp procedure that fails on purpose. Takes a second or two, and resets itself between scenarios so you can watch all three run back to back.\nSET NOCOUNT ON; CREATE OR ALTER PROCEDURE dbo.sp_ResetLedgerDemo AS BEGIN IF OBJECT_ID(\u0026#39;tempdb..#Ledger\u0026#39;) IS NOT NULL DROP TABLE #Ledger; CREATE TABLE #Ledger (Seq INT IDENTITY, Step varchar(300)); IF OBJECT_ID(\u0026#39;tempdb..#ChildStep\u0026#39;) IS NOT NULL DROP PROCEDURE #ChildStep; END GO -- (If your client doesn\u0026#39;t support GO as a batch separator, run the two -- halves of this script as separate executions instead.) EXEC dbo.sp_ResetLedgerDemo; CREATE PROCEDURE #ChildStep AS BEGIN INSERT INTO #Ledger (Step) VALUES (\u0026#39;child: before error\u0026#39;); SELECT 1/0; -- deliberate error INSERT INTO #Ledger (Step) VALUES (\u0026#39;child: after error\u0026#39;); END; PRINT \u0026#39;=== Scenario 1: XACT_ABORT OFF, no TRY/CATCH ===\u0026#39;; SET XACT_ABORT OFF; BEGIN TRAN; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: before EXEC\u0026#39;); EXEC #ChildStep; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: after EXEC\u0026#39;); COMMIT; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: COMMIT succeeded\u0026#39;); SELECT Step FROM #Ledger ORDER BY Seq; IF @@TRANCOUNT \u0026gt; 0 ROLLBACK; EXEC dbo.sp_ResetLedgerDemo; CREATE PROCEDURE #ChildStep AS BEGIN INSERT INTO #Ledger (Step) VALUES (\u0026#39;child: before error\u0026#39;); SELECT 1/0; INSERT INTO #Ledger (Step) VALUES (\u0026#39;child: after error\u0026#39;); END; PRINT \u0026#39;=== Scenario 2: TRY/CATCH added, XACT_ABORT still OFF ===\u0026#39;; SET XACT_ABORT OFF; BEGIN TRY BEGIN TRAN; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: before EXEC\u0026#39;); EXEC #ChildStep; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: after EXEC\u0026#39;); COMMIT; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: COMMIT succeeded\u0026#39;); END TRY BEGIN CATCH -- Note the ordering here is the point of the demo: this logs BEFORE -- rolling back, so watch what survives. INSERT INTO #Ledger (Step) VALUES (\u0026#39;CATCH fired, XACT_STATE=\u0026#39; + CAST(XACT_STATE() AS varchar(3))); IF XACT_STATE() \u0026lt;\u0026gt; 0 ROLLBACK; INSERT INTO #Ledger (Step) VALUES (\u0026#39;CATCH: after rollback check\u0026#39;); END CATCH; SELECT Step FROM #Ledger ORDER BY Seq; IF @@TRANCOUNT \u0026gt; 0 ROLLBACK; EXEC dbo.sp_ResetLedgerDemo; CREATE PROCEDURE #ChildStep AS BEGIN INSERT INTO #Ledger (Step) VALUES (\u0026#39;child: before error\u0026#39;); SELECT 1/0; INSERT INTO #Ledger (Step) VALUES (\u0026#39;child: after error\u0026#39;); END; PRINT \u0026#39;=== Scenario 3: XACT_ABORT ON, rollback FIRST in CATCH ===\u0026#39;; SET XACT_ABORT ON; BEGIN TRY BEGIN TRAN; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: before EXEC\u0026#39;); EXEC #ChildStep; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: after EXEC\u0026#39;); COMMIT; INSERT INTO #Ledger (Step) VALUES (\u0026#39;outer: COMMIT succeeded\u0026#39;); END TRY BEGIN CATCH DECLARE @state INT = XACT_STATE(), @msg varchar(300) = ERROR_MESSAGE(), @proc varchar(300) = ISNULL(ERROR_PROCEDURE(), \u0026#39;(top level)\u0026#39;), @line INT = ERROR_LINE(); IF XACT_STATE() \u0026lt;\u0026gt; 0 ROLLBACK; -- always first INSERT INTO #Ledger (Step) VALUES (\u0026#39;CATCH: state=\u0026#39; + CAST(@state AS varchar(3)) + \u0026#39; proc=\u0026#39; + @proc + \u0026#39; line=\u0026#39; + CAST(@line AS varchar(10)) + \u0026#39; msg=\u0026#39; + @msg); END CATCH; SELECT Step FROM #Ledger ORDER BY Seq; IF @@TRANCOUNT \u0026gt; 0 ROLLBACK; SET XACT_ABORT OFF; DROP PROCEDURE IF EXISTS dbo.sp_ResetLedgerDemo; Requires SQL Server 2016+. Everything lives in temp objects that vanish when your session ends, safe to run anywhere.\nThe lesson The default error-handling behavior in SQL Server is \u0026ldquo;keep going,\u0026rdquo; not \u0026ldquo;stop.\u0026rdquo; If you want a multi-step script to either fully happen or fully not happen, you have to ask for that explicitly with SET XACT_ABORT ON, and get the order of operations in your CATCH block right, or your own diagnostics can vanish along with everything else.\nIf you\u0026rsquo;ve got a script that calls other scripts without SET XACT_ABORT ON, or a CATCH block that logs before it rolls back, it\u0026rsquo;s worth five minutes to go check which of these you\u0026rsquo;re sitting on.\n","permalink":"https://hanhpham.vercel.app/posts/sql-server-trustworthy-rollback-xact-abort/","summary":"When one step in a multi-step SQL Server script fails, what actually happens to your data depends entirely on how you\u0026rsquo;ve wired up error handling, and the default is not what most people assume. Three scenarios, measured with real ledger output.","title":"Making a Multi-Step SQL Server Script's Rollback Actually Trustworthy"},{"content":"Picture two tables. One lists records to evaluate: accounts, orders, whatever your domain is. The other lists rules, and a rule is allowed to leave a field blank, meaning \u0026ldquo;this applies to everyone\u0026rdquo; rather than one specific value. That\u0026rsquo;s a completely reasonable design. Here\u0026rsquo;s the join that reads it:\n-- a blank field on the rules side means \u0026#34;match anything here\u0026#34; INSERT INTO #AccountCondition SELECT ... FROM #Accounts A INNER JOIN #Rules R ON (A.RegionCode = R.RegionCode OR R.RegionCode IS NULL) AND (A.SegmentCode = R.SegmentCode OR R.SegmentCode IS NULL) Each \u0026ldquo;match this, or leave it blank\u0026rdquo; condition makes sense on its own. But once a meaningful share of the rules leave a field blank, each of those rules matches every account in the batch: not one, but all of them. This isn\u0026rsquo;t a hypothetical. I ran it:\nInput Result 20,000 accounts, 4,000 rules (~28% of rules leave a field blank) n/a Rows produced by the join above 646,400 32 rows per account, from a join that reads like an ordinary filter. Nobody wrote a bug: every individual condition is correct. The multiplication is emergent: a cartesian-style blowup that doesn\u0026rsquo;t look like a mistake, it looks like a normal join, right up until you count the rows.\n646,400 rows alone usually isn\u0026rsquo;t fatal; that\u0026rsquo;s nothing for SQL Server. The real damage happens once that oversized result feeds into cleanup logic downstream.\nGive your staging table an index, then check what it actually bought you A common next step: clean up that oversized result by deleting rows that don\u0026rsquo;t belong, checked against a lookup table.\nDELETE tgt FROM #AccountCondition tgt -- no index: a heap LEFT JOIN ScopeHierarchy H ON ISNULL(tgt.ScopeCode, \u0026#39;*\u0026#39;) = H.ScopeCode AND (tgt.SegmentCode = H.Level1 OR tgt.SegmentCode = H.Level2 OR ... -- through Level10 ) WHERE H.ScopeCode IS NULL The table holding those rows is a heap: no sort order, so finding anything means checking every row. The obvious fix is an index:\nCREATE CLUSTERED INDEX CIX ON #AccountCondition (ScopeCode, SegmentCode); Here\u0026rsquo;s the part worth actually checking, rather than assuming: does that index get used the way you\u0026rsquo;d expect? I ran the delete above, indexed and un-indexed, and pulled the real execution plan for each.\nVersion Time vs. baseline Plan shows No index, OR chain (baseline) 1,614.6 ms 1.00× Table Scan + Index only, same OR chain 852.4 ms 1.89× Clustered Index Scan Measured on 100,000 accounts, 20,000 rules, a 1,000-row hierarchy table: 16.16 million rows entering the delete.\nThe index made things faster, but the plan confirms it\u0026rsquo;s still scanning every row front to back, not seeking. That gap is the whole point of this post.\nWrapping a comparison in ISNULL() blocks SQL Server from using an index to jump straight to matching rows; it has to evaluate the function against every row instead. This is called non-sargable, and it\u0026rsquo;s invisible unless you actually read the execution plan. A CREATE INDEX line is not proof that the index is being used the way you think.\nThe ~1.9× gain here came from the index reorganizing physical storage on disk, a smaller, roughly constant win. The bigger prize, an actual seek, needs one more change.\nTen-way OR chains vs. UNPIVOT The other half of that query checks one column against ten separate lookup columns, because the hierarchy table was modeled wide instead of tall. UNPIVOT turns that around: \u0026ldquo;one row, ten columns to compare\u0026rdquo; becomes \u0026ldquo;up to ten rows, one column, seekable\u0026rdquo;:\n;WITH ScopeLevels AS ( SELECT ScopeCode, LevelValue FROM ScopeHierarchy UNPIVOT (LevelValue FOR LevelCol IN (Level1, Level2, ..., Level10)) u ) DELETE tgt FROM #AccountCondition tgt WHERE tgt.SegmentCode IS NOT NULL AND NOT EXISTS ( SELECT 1 FROM ScopeLevels s WHERE s.ScopeCode = ISNULL(tgt.ScopeCode, \u0026#39;*\u0026#39;) AND s.LevelValue = tgt.SegmentCode ); I measured all four combinations on the same 16.16M-row input, not just the two ends:\nVersion Time vs. baseline A: heap, OR chain, ISNULL (baseline) 1,614.6 ms 1.00× B: indexed, OR chain, ISNULL 852.4 ms 1.89× C: heap, UNPIVOT 1,040.2 ms 1.55× D: indexed, UNPIVOT (both fixes) 909.1 ms 1.78× All four deleted the exact same rows, verified by comparing the remaining row count afterward.\nNotice B and D land close enough together that a single run each doesn\u0026rsquo;t support a clean ranking between them: that\u0026rsquo;s noise, not a result, and I\u0026rsquo;m not going to pretend otherwise. What the execution plan does confirm is the part that matters: only the indexed-plus-UNPIVOT version is structurally capable of a true index seek. A scan-based win like B is roughly proportional to row count and flattens out; a seek keeps paying off disproportionately as the table grows, which single-run numbers at one fixed size can\u0026rsquo;t show you by themselves. If this table is 16 million rows today and 160 million next year, B and D stop being close.\nTwo more changes are worth doing even though I didn\u0026rsquo;t isolate their timing impact in this test: separate the wildcard rules from the specific rules before the original join, so a \u0026ldquo;match anything\u0026rdquo; rule isn\u0026rsquo;t physically copied once per matching account. And filter the input down to what the caller actually needs before handing it downstream: remove rows before the multiplication happens, rather than cleaning them up after.\nTry it yourself This script is self-contained: everything lives in temp tables that vanish when your session ends. It builds the same 100,000-account / 20,000-rule / 1,000-row-hierarchy dataset used above and runs all four delete variants back to back. Takes under a minute, most of it spent generating the data, not running the deletes.\nSET NOCOUNT ON; DROP TABLE IF EXISTS #Accounts, #Rules, #ScopeHierarchy, #P_Baseline, #P_IndexOnly, #P_UnpivotOnly, #P_Both; -- 100,000 accounts, 20,000 rules (70% blank segment, 10% blank region, -- 20% both set), a 1,000-row hierarchy table with 10 \u0026#34;level\u0026#34; columns. ;WITH n AS ( SELECT TOP (100000) ROW_NUMBER() OVER (ORDER BY (SELECT NULL)) - 1 AS n FROM sys.all_objects a CROSS JOIN sys.all_objects b CROSS JOIN sys.all_objects c ) SELECT n AS RowNum, \u0026#39;R\u0026#39; + CAST(n % 100 AS varchar(5)) AS RegionCode, \u0026#39;S\u0026#39; + CAST((n * 7) % 100 AS varchar(5)) AS SegmentCode INTO #Accounts FROM n; ;WITH n AS ( SELECT TOP (20000) ROW_NUMBER() OVER (ORDER BY (SELECT NULL)) - 1 AS n FROM sys.all_objects a CROSS JOIN sys.all_objects b CROSS JOIN sys.all_objects c ) SELECT n AS RuleID, CASE WHEN n \u0026gt;= 14000 AND n \u0026lt; 16000 THEN NULL ELSE \u0026#39;R\u0026#39; + CAST(n % 100 AS varchar(5)) END AS RegionCode, CASE WHEN n \u0026lt; 14000 THEN NULL ELSE \u0026#39;S\u0026#39; + CAST((n * 3) % 100 AS varchar(5)) END AS SegmentCode INTO #Rules FROM n; ;WITH n AS ( SELECT TOP (1000) ROW_NUMBER() OVER (ORDER BY (SELECT NULL)) - 1 AS n FROM sys.all_objects a CROSS JOIN sys.all_objects b ) SELECT n AS HierarchyID, \u0026#39;R\u0026#39; + CAST(n % 100 AS varchar(5)) AS ScopeCode, \u0026#39;S\u0026#39; + CAST(n % 100 AS varchar(5)) AS Level1, \u0026#39;S\u0026#39; + CAST((n*3)%100 AS varchar(5)) AS Level2, \u0026#39;S\u0026#39; + CAST((n*7)%100 AS varchar(5)) AS Level3, \u0026#39;S\u0026#39; + CAST((n*9)%100 AS varchar(5)) AS Level4, \u0026#39;S\u0026#39; + CAST((n*11)%100 AS varchar(5)) AS Level5, \u0026#39;S\u0026#39; + CAST((n*13)%100 AS varchar(5)) AS Level6, \u0026#39;S\u0026#39; + CAST((n*17)%100 AS varchar(5)) AS Level7, \u0026#39;S\u0026#39; + CAST((n*19)%100 AS varchar(5)) AS Level8, \u0026#39;S\u0026#39; + CAST((n*21)%100 AS varchar(5)) AS Level9, \u0026#39;S\u0026#39; + CAST((n*23)%100 AS varchar(5)) AS Level10 INTO #ScopeHierarchy FROM n; SELECT A.RowNum, A.RegionCode AS ScopeCode, A.SegmentCode INTO #P_Baseline FROM #Accounts A INNER JOIN #Rules R ON (A.RegionCode = R.RegionCode OR R.RegionCode IS NULL) AND (A.SegmentCode = R.SegmentCode OR R.SegmentCode IS NULL); PRINT \u0026#39;Rows entering the delete: \u0026#39; + CAST(@@ROWCOUNT AS varchar(20)); SELECT * INTO #P_IndexOnly FROM #P_Baseline; SELECT * INTO #P_UnpivotOnly FROM #P_Baseline; SELECT * INTO #P_Both FROM #P_Baseline; PRINT \u0026#39;--- A) heap, OR chain, ISNULL (baseline) ---\u0026#39;; SET STATISTICS TIME ON; DELETE tgt FROM #P_Baseline tgt LEFT JOIN #ScopeHierarchy H ON ISNULL(tgt.ScopeCode, \u0026#39;*\u0026#39;) = H.ScopeCode AND (tgt.SegmentCode = H.Level1 OR tgt.SegmentCode = H.Level2 OR tgt.SegmentCode = H.Level3 OR tgt.SegmentCode = H.Level4 OR tgt.SegmentCode = H.Level5 OR tgt.SegmentCode = H.Level6 OR tgt.SegmentCode = H.Level7 OR tgt.SegmentCode = H.Level8 OR tgt.SegmentCode = H.Level9 OR tgt.SegmentCode = H.Level10) WHERE H.ScopeCode IS NULL; SET STATISTICS TIME OFF; PRINT \u0026#39;--- B) indexed, OR chain, ISNULL (index alone) ---\u0026#39;; CREATE CLUSTERED INDEX CIX ON #P_IndexOnly (ScopeCode, SegmentCode); SET STATISTICS TIME ON; DELETE tgt FROM #P_IndexOnly tgt LEFT JOIN #ScopeHierarchy H ON ISNULL(tgt.ScopeCode, \u0026#39;*\u0026#39;) = H.ScopeCode AND (tgt.SegmentCode = H.Level1 OR tgt.SegmentCode = H.Level2 OR tgt.SegmentCode = H.Level3 OR tgt.SegmentCode = H.Level4 OR tgt.SegmentCode = H.Level5 OR tgt.SegmentCode = H.Level6 OR tgt.SegmentCode = H.Level7 OR tgt.SegmentCode = H.Level8 OR tgt.SegmentCode = H.Level9 OR tgt.SegmentCode = H.Level10) WHERE H.ScopeCode IS NULL; SET STATISTICS TIME OFF; PRINT \u0026#39;--- C) heap, UNPIVOT (unpivot alone) ---\u0026#39;; SET STATISTICS TIME ON; ;WITH ScopeLevels AS ( SELECT ScopeCode, LevelValue FROM #ScopeHierarchy UNPIVOT (LevelValue FOR LevelCol IN (Level1,Level2,Level3,Level4,Level5,Level6,Level7,Level8,Level9,Level10)) AS u ) DELETE tgt FROM #P_UnpivotOnly tgt WHERE tgt.SegmentCode IS NOT NULL AND NOT EXISTS (SELECT 1 FROM ScopeLevels s WHERE s.ScopeCode = ISNULL(tgt.ScopeCode, \u0026#39;*\u0026#39;) AND s.LevelValue = tgt.SegmentCode); SET STATISTICS TIME OFF; PRINT \u0026#39;--- D) indexed, UNPIVOT (both fixes) ---\u0026#39;; CREATE CLUSTERED INDEX CIX ON #P_Both (ScopeCode, SegmentCode); SET STATISTICS TIME ON; ;WITH ScopeLevels AS ( SELECT ScopeCode, LevelValue FROM #ScopeHierarchy UNPIVOT (LevelValue FOR LevelCol IN (Level1,Level2,Level3,Level4,Level5,Level6,Level7,Level8,Level9,Level10)) AS u ) DELETE tgt FROM #P_Both tgt WHERE tgt.SegmentCode IS NOT NULL AND NOT EXISTS (SELECT 1 FROM ScopeLevels s WHERE s.ScopeCode = ISNULL(tgt.ScopeCode, \u0026#39;*\u0026#39;) AND s.LevelValue = tgt.SegmentCode); SET STATISTICS TIME OFF; SELECT \u0026#39;A baseline\u0026#39; AS Version, COUNT(*) AS RowsRemaining FROM #P_Baseline UNION ALL SELECT \u0026#39;B index-only\u0026#39;, COUNT(*) FROM #P_IndexOnly UNION ALL SELECT \u0026#39;C unpivot-only\u0026#39;, COUNT(*) FROM #P_UnpivotOnly UNION ALL SELECT \u0026#39;D both\u0026#39;, COUNT(*) FROM #P_Both; -- All four RowsRemaining values should match exactly. Requires SQL Server 2016+. Every object is a #temp table or CTE; nothing touches a real database, safe to run anywhere.\nThe lesson A wildcard rule multiplies exactly, not gently, once both the rules table and the data table grow, so know your blank-field ratio before you trust a row estimate. And an index can help even when it can\u0026rsquo;t be seeked, but don\u0026rsquo;t stop there: wrapping a join column in a function blocks true seeking no matter what indexes exist. Fix the predicate and the index together, and check the actual execution plan rather than assuming a CREATE INDEX line settled it.\nThat still leaves one more failure mode in this neighborhood: what actually happens to a multi-step script\u0026rsquo;s data when one step fails partway through, covered in the next post.\n","permalink":"https://hanhpham.vercel.app/posts/sql-server-wildcard-joins-and-unpivot/","summary":"A join rule that matches \u0026lsquo;anything\u0026rsquo; looks like an ordinary filter, until it quietly multiplies your row count. Measured, not estimated: what a wildcard join actually costs, why a CREATE INDEX line doesn\u0026rsquo;t always buy you a seek, and why UNPIVOT beats a ten-way OR chain.","title":"When a Wildcard Join Multiplies Your Data (and How to Actually Fix It)"},{"content":"Most Go releases are easy to summarize: a few library additions, a compiler improvement, a port dropped. You skim the notes, note the one thing relevant to you, and move on.\nGo 1.27, released on 19 August 2026, is not that. Several changes in this release alter how your program behaves without you writing a single new line of code: the JSON package you already import, the allocator underneath every make, the HTTP/2 server you already run, and the vet checks your CI already invokes. That combination is unusual, and it\u0026rsquo;s why this one deserves more than a skim.\nHere\u0026rsquo;s what\u0026rsquo;s actually in it, and what it will cost you.\nThe short version If you read nothing else:\nGeneric methods landed. Methods can now declare their own type parameters. This closes the most-cited hole left open when generics shipped in 1.18. encoding/json/v2 exists, and you\u0026rsquo;re already running it. The v1 package is now a compatibility layer over v2. You don\u0026rsquo;t have to migrate. You are nonetheless on new code. UUIDs are in the standard library. One of the most reflexively-added dependencies in the ecosystem just became an import. Small allocations got up to 30% faster, on by default, costing a flat ~60 KB of binary. goroutineleak profiling is generally available, and it works by reusing the garbage collector\u0026rsquo;s reachability analysis, which is genuinely clever. Several changes will break something in your CI. They\u0026rsquo;re at the bottom of this post. Read that section before you upgrade, not after. Generic methods, and the two restrictions that define them Since 1.18, a function could be generic but a method could only use the type parameters of its receiver. If you wanted a method parameterized independently of its type, you wrote a free function and apologized about it.\nThat\u0026rsquo;s over:\n// a method declaring its own type parameter, new in 1.27 func (r *Rand) N[Int intType](n Int) Int The standard library uses it immediately: math/rand/v2 gains a generic Rand.N matching the behavior of the existing top-level N.\nBut the interesting part is what you still can\u0026rsquo;t do:\nInterface methods cannot declare type parameters. A generic method cannot implement an interface method. Read those together and the design becomes clear. Generic methods are a facility for concrete types. The interface system stays monomorphic, which is exactly what keeps Go\u0026rsquo;s method dispatch a single indirect call instead of a runtime instantiation problem.\nIf you were hoping generic methods would let you write a generic interface, that\u0026rsquo;s not what shipped, and it isn\u0026rsquo;t an oversight. It\u0026rsquo;s the price of dispatch staying cheap.\nTwo smaller language changes ride along. Struct literal keys may now be any valid field selector, not just top-level field names, so you can initialize a nested field directly in a composite literal. And function type inference is generalized to apply everywhere a generic function is assigned to a variable or converted to a matching function type, rather than only in call position. Neither will change your architecture; both will delete a few lines.\njson/v2: the migration you don\u0026rsquo;t have to do encoding/json/v2 is a full revision of Go\u0026rsquo;s most-used package. Six entry points (Marshal, MarshalWrite, MarshalEncode, Unmarshal, UnmarshalRead, UnmarshalDecode), all taking variadic Options. Unmarshaling is significantly faster than v1; marshaling is broadly at parity.\nIt\u0026rsquo;s also stricter by default. v2 rejects two things v1 quietly accepted:\ninvalid UTF-8 in JSON strings duplicate names within a JSON object Both of those are, correctly, errors. Both of them almost certainly exist somewhere in the data your service receives today.\nNow the part worth pausing on: the original encoding/json is now backed by v2, with options configured to preserve v1 semantics. Your v1 code keeps working. No migration is required. There\u0026rsquo;s an escape hatch (GOEXPERIMENT=nojsonv2) that\u0026rsquo;s expected to be removed in a future release.\nYour v1 code keeps its semantics, but not necessarily its exact error strings; the release notes call this out explicitly. If you have tests asserting on JSON error text, or an API that forwards the unmarshal error verbatim to a client, that\u0026rsquo;s where this surfaces.\nSo the honest framing is: you did not choose to adopt json/v2, but you are running it. The compatibility layer is well-tested and the Go team has done this kind of thing carefully before. Still, if you have a service whose entire job is chewing JSON, this is the release where you benchmark before and after rather than assuming.\nThere\u0026rsquo;s also encoding/json/jsontext for the layer below: an Encoder and Decoder over a stream of Token and Value, with the validity state machine handled for you. If you\u0026rsquo;ve ever hand-rolled a streaming JSON scanner, this is that, done properly.\nuuid in the standard library Small change. Wide blast radius.\nGenerating and parsing UUIDs is now stdlib. That\u0026rsquo;s it. But \u0026ldquo;add a UUID library\u0026rdquo; is one of the most automatic dependency decisions in Go, and it just stopped being a decision. Expect this to quietly drop a line from a very large number of go.mod files over the next year.\nThe allocator, and a rare shape of trade-off The compiler now emits size-specialized allocation routines. The numbers:\nAllocations under 80 bytes up to 30% faster Allocation-heavy programs, overall ~1% Binary size cost ~60 KB, fixed That last row is what makes this interesting. The cost is a constant, independent of workload. The benefit scales with how much you allocate. There aren\u0026rsquo;t many optimizations shaped like that, and it means the trade is close to strictly good for any server-side binary: 60 KB is noise against a Go binary, and 1% of CPU across a fleet is not.\nIt\u0026rsquo;s on by default. GOEXPERIMENT=nosizespecializedmalloc opts out and is expected to disappear in Go 1.28, so if you find yourself reaching for it, treat that as a bug report to file, not a setting to keep.\nGoroutine leak detection, and its blind spot The goroutineleak profile graduates from experiment to generally available, exposed through runtime/pprof and /debug/pprof/goroutineleak:\ngo tool pprof http://localhost:6060/debug/pprof/goroutineleak It reports goroutines blocked on a concurrency primitive (a channel, a sync.Mutex, a sync.Cond) that cannot be unblocked. The detection method is the nice part: it uses the garbage collector\u0026rsquo;s existing reachability analysis. If nothing can reach the primitive a goroutine is parked on, nothing can ever wake it. That\u0026rsquo;s a leak, provably, and the GC was already computing the reachability graph for other reasons.\nThe blind spot is stated plainly in the release notes and you should hold onto it: it cannot detect leaks blocked on primitives reachable through global variables, or through the locals of a runnable goroutine. In both cases the primitive is still reachable, so the analysis can\u0026rsquo;t prove nothing will signal it.\nPractically: an empty profile is not proof of no leaks. It\u0026rsquo;s proof of no provable leaks. That\u0026rsquo;s still a large improvement over what you had, which was reading goroutine dumps by hand at 2am.\nPost-quantum signatures New crypto/mldsa implements ML-DSA (FIPS 204). crypto/x509 handles ML-DSA keys and signatures. crypto/tls negotiates them in TLS 1.3 through MLDSA44, MLDSA65 and MLDSA87, and adds MLKEM1024 key exchange.\nGo continues to be unusually early on post-quantum crypto. If you have a harvest-now-decrypt-later threat model, the signature half of that story is now available without a third-party dependency and without writing any of it yourself.\nWhat will actually cost you time Go keeps the Go 1 compatibility promise, so almost everything still compiles and runs. But a handful of changes produce failures that won\u0026rsquo;t look like upgrade problems, which is what makes them expensive.\n1. compress/flate output bytes changed Compression got faster, and the exact encoded output differs from 1.26. This ripples through archive/zip, compress/gzip, compress/zlib and image/png.\nIf you have golden-file tests, artifact checksums, or reproducible-build assertions over compressed output, they break. And the failure presents as \u0026ldquo;our PNG encoder is producing wrong bytes\u0026rdquo;, which is a bad thing to start debugging at the wrong end.\nThis is the one I\u0026rsquo;d check first.\n2. Function literal names changed Closures now get simpler generated names, identical whether or not they\u0026rsquo;re inlined, and identical instances may share code in the binary.\nTwo consequences. The obvious one: tests asserting on symbol names fail. The subtle one, from the release notes: this may expose incorrect function code pointer equality comparisons, because two closures with different captured data can now end up with equal code pointers. If any of your code compares function pointers for equality, it was already relying on undefined behavior, and 1.27 is where you find out.\n3. go test now runs the stdversion vet check by default It reports uses of standard library symbols newer than the Go version in force for that file (determined by the go directive in go.mod plus build tags). That\u0026rsquo;s a real correctness win: it catches the class of bug where your code compiles locally and fails on an older toolchain.\nIt will also turn some currently-green builds red on first upgrade. Budget for that rather than being surprised by it.\n4. asynctimerchan is gone The GODEBUG added in 1.23 has been removed. Channels created by the time package are now always synchronous. If you set this to 1 to preserve pre-1.23 buffering behavior, that door is closed and the code depending on it has to change.\n5. HTTP/2 now honours client priority signals The HTTP/2 server implements RFC 9218 and will prioritize streams the client marks as higher priority, instead of the previous round-robin. For most services this is an improvement you\u0026rsquo;d have asked for. But it is a change in which response finishes first under concurrent load, arriving by default, and if you have latency assertions or a load test tuned against round-robin behavior, it will move your numbers. Server.DisableClientPriority = true restores the old scheduling.\n6. Unicode 15 → 17 The unicode tables jump two major versions. IsLetter, IsDigit, category lookups, and anything built on them can return different answers for characters that were unassigned before. If you validate user input against Unicode categories, your accept/reject boundary moved.\nAlso worth a scan macOS 12 support ended. Darwin requires macOS 13 Ventura or later. Check CI runners before laptops. Five TLS/x509 GODEBUGs removed permanently: tlsunsafeekm, tlsrsakex, tls3des, tls10server and x509keypairleaf. These were the compatibility ramps off deprecated crypto; the ramps are gone. (tlskyber also shows up in the notes, but it was actually removed back in 1.24 and is only now documented as removed.) gotypesalias is gone too: go/types now always produces an Alias node, which matters if you write analysis tooling. go tool trace -http=:6060 now binds localhost only when given just a port, matching go tool pprof. If you profile on a remote box and connect from your laptop, this looks like the tool silently not working. Pass an explicit -http=0.0.0.0:6060. go mod tidy restructures your require blocks for modules declaring go 1.27, with duplicates merged, collapsed to at most two blocks (direct and indirect). Correct, and a large one-time diff nobody scheduled. Goroutine labels now appear in traceback headers for go 1.27 modules. Better panics; also a changed log format if something downstream parses them. GODEBUG=tracebacklabels=0 opts out, and that opt-out is expected to stay indefinitely. HTTP/1 response bodies now auto-drain on close, up to a conservative limit. Better connection reuse in general; a performance regression in rare misconfigured cases. Escape via Transport.DisableKeepAlives = true. SystemCertPool honours SSL_CERT_FILE and SSL_CERT_DIR on Windows and Darwin, and when set, Go uses its own verifier rather than the platform APIs. That\u0026rsquo;s a change in which verifier runs, not just which roots load. Opt out with GODEBUG=x509sslcertoverrideplatform=0. net.UnixConn read methods return io.EOF directly instead of wrapping it in a net.OpError. Code doing a type assertion to *net.OpError on that path stops matching. crypto/ecdsa\u0026rsquo;s PrivateKey.Sign now validates hash length when given non-nil SignerOpts. A latent mismatch that used to sign anyway now errors. ppc64 big-endian moved to the ELFv2 ABI, requiring kernel 3.13+, and gaining cgo, PIE and external linking in exchange. Bazaar (bzr) support removed from the go command. Three tiers of commitment The release notes present everything uniformly, but Go 1.27 actually ships at three quite different levels of promise. This distinction matters when you\u0026rsquo;re deciding what\u0026rsquo;s allowed into production code.\nStable: covered by the Go 1 compatibility promise. Generic methods, struct-literal field selectors, generalized type inference, encoding/json/v2, jsontext, crypto/mldsa, uuid, the goroutineleak profile, and every minor library addition (strings.CutLast, bytes.CutLast, URL.Clone, Values.Clone, maphash.Hasher, math/big\u0026rsquo;s Int.Divide, httptest.NewTestServer, synctest.Sleep, and the rest). Use freely.\nOn by default, with an expiring opt-out. Size-specialized allocation (nosizespecializedmalloc, expected gone in 1.28) and json/v2 backing v1 (nojsonv2, removal TBD). These are not settings. They\u0026rsquo;re stopgaps with a deadline, and treating them as configuration is how you end up blocked on a toolchain upgrade next year.\nExperimental: behind a build flag, API explicitly unstable. The new portable simd package and the architecture-specific simd/archsimd, both behind GOEXPERIMENT=simd. archsimd revises the amd64 API and adds arm64 Neon and WebAssembly support: 128-bit vectors on wasm/arm64/amd64, with 256- and 512-bit on some amd64 parts. Fascinating, and spike-only.\nAn upgrade checklist Grep for golden-file or checksum tests over zip/gzip/zlib/png output. Fix those first. Run go vet before you run go test, and see what stdversion says while it\u0026rsquo;s still advisory in your head. Search for function-pointer equality comparisons. If you find any, they were already broken. Confirm CI runners are on macOS 13+. Check whether you set asynctimerchan or any of the removed TLS/x509 GODEBUGs. If you have HTTP/2 latency assertions or Unicode-category input validation, re-run those suites specifically. If JSON throughput is on your critical path, benchmark before and after. You\u0026rsquo;re on new code whether you migrated or not. Then enjoy the free 1%. The read Go 1.27\u0026rsquo;s headline is generic methods, and that\u0026rsquo;s a real, long-awaited language change. But the headline isn\u0026rsquo;t where the release earns its keep.\nThe pattern worth noticing is that most of the consequential changes (json/v2 under v1, the allocator, the stdversion vet check, HTTP/2 prioritization, the Unicode bump, traceback labels) take effect without you opting in. That\u0026rsquo;s a lot of behavioral change to absorb through a version bump, and it\u0026rsquo;s a mild departure from the conservatism Go usually shows here.\nIt\u0026rsquo;s all defensible. The compatibility layer is careful, the allocator trade is close to free, and the vet check catches real bugs. But it means this is a release to upgrade deliberately (read the notes, run the checklist, then move) rather than one to bump on a Friday and see what happens.\nSources: Go 1.27 Release Notes · Go 1.27 is released · The Go Blog, 19 August 2026\n","permalink":"https://hanhpham.vercel.app/posts/go-1-27-is-quietly-load-bearing/","summary":"Generic methods finally landed, json/v2 is already running underneath your code, and the allocator got faster while you weren\u0026rsquo;t looking. Six of this release\u0026rsquo;s changes take effect without you opting in; here\u0026rsquo;s which ones will cost you an afternoon.","title":"Go 1.27 Is Quietly Load-Bearing"},{"content":"We\u0026rsquo;ve spent five posts building up a toolkit: boundaries via coupling and cohesion, communication styles, the technology to implement them, and sagas for cross-service consistency. Each one made sense on its own. The real test is whether they compose, so let\u0026rsquo;s design one actual workflow from scratch and make every decision in order: MusicCorp shipping a CD order.\nStep 0: what are the moving pieces? Before any communication decision, boundaries first. Following the coupling analysis from post one, MusicCorp\u0026rsquo;s order flow involves five services, each owning one aggregate and its state machine:\nOrder: owns the order\u0026rsquo;s lifecycle (PLACED → PAID → PICKING → SHIPPED → COMPLETED), and is the only thing allowed to decide whether a requested transition is valid. This is the fix for the common-coupling trap from post one; nobody else gets to mutate order status directly. Payment: takes payment, knows nothing about warehouses or shipping. Warehouse: reserves stock, packages, and dispatches. Loyalty: awards points. Notifications: emails the customer at various points. Each is cohesive (their own concern lives entirely inside their own boundary) and only domain-coupled to the others where genuinely necessary.\nStep 1: pick the communication style, service by service Now we ask, per interaction, the two questions from post two: does the caller block, and is this a directed request or a broadcast? Not every hop gets the same answer.\nPlacing an order and taking payment is a case where the caller genuinely needs to know the result before continuing: a customer waiting on checkout can\u0026rsquo;t be told \u0026ldquo;we\u0026rsquo;ll email you in a few hours about whether your card worked.\u0026rdquo; This is synchronous request-response: Order calls Payment, blocks, and gets a definite yes or no back before the checkout page can respond.\nReserving stock is a request-response too, since checkout has to know whether the item is even available, but it doesn\u0026rsquo;t strictly need to block the customer for the entire duration if the warehouse system is slow, so an asynchronous request-response over a queue is a reasonable choice here without changing the shape of the interaction.\nPackaging, dispatch, awarding points, and sending notifications are a different story. None of these needs to happen inside the request that the customer is waiting on; packaging can take hours or days, and there\u0026rsquo;s no reason Order should know or care that Loyalty and Notifications exist at all. This is where event-driven collaboration earns its complexity: Warehouse fires a Stock Reserved event, Payment fires a Payment Taken event, and both Loyalty and Warehouse react to Payment Taken independently and in parallel: one awards points, the other dispatches the package. Neither service told the other to do anything; they just each reacted to a fact that was broadcast.\nOrderProcessor Payment Warehouse | | | |--- reserve stock ---\u0026gt;| | | | | |\u0026lt;---- reserved -------| | | | | |--- take payment ----\u0026gt;| | | | | | [Payment Taken event fires] | | | | | (Loyalty reacts) (Warehouse reacts: | awards points dispatch package) Notice this is exactly the \u0026ldquo;mix and match\u0026rdquo; point from post two: the same workflow uses synchronous request-response where an answer is genuinely needed right away, and event-driven collaboration everywhere it isn\u0026rsquo;t. Neither choice is \u0026ldquo;more correct\u0026rdquo; in general; each fits its specific hop.\nStep 2: pick the technology per interaction With styles chosen, the technology conversation from post three gets short. Order-to-Payment, being synchronous request-response with a small number of well-controlled internal consumers, is a natural fit for gRPC: good performance, strong schemas, and MusicCorp controls both ends. The event-driven hops (Payment Taken, Stock Reserved) need a topic, not a queue, since multiple independent consumer groups (Loyalty, Warehouse, potentially a future Recommendations service) each need their own copy of the same event; this is where a broker like Kafka or a managed equivalent earns its keep, particularly for the message permanence that lets a newly deployed consumer catch up on history it missed.\nWhatever\u0026rsquo;s chosen, the events themselves should be fully detailed rather than \u0026ldquo;just an ID\u0026rdquo;: Notifications needs a name and email address to send a personalized message, and making it call back to Order for that information every time would reintroduce exactly the domain coupling event-driven collaboration was supposed to avoid.\nStep 3: what happens when something fails partway through? This is where the saga thinking from post four becomes unavoidable. The order fulfillment process spans five services and cannot be one ACID transaction, so what happens if packaging fails because the CD isn\u0026rsquo;t actually on the shelf, despite the system thinking it was?\nThis is a choreographed saga: no single orchestrator, each service reacting to events and deciding its own next move. Consider the failure case explicitly:\nOrder Placed -\u0026gt; Stock Reserved -\u0026gt; Payment Taken -\u0026gt; [Packaging fails: item missing] By this point, payment has already been taken and (if we hadn\u0026rsquo;t applied the reordering trick from post four) loyalty points may already have been awarded. Rolling the whole order back now means firing compensating transactions: refund the payment, and reverse the loyalty award if it already happened. Neither of these is a true rollback: refunding isn\u0026rsquo;t \u0026ldquo;pretend the charge never happened,\u0026rdquo; it\u0026rsquo;s a new transaction that reverses the effect. If a \u0026ldquo;sorry, your order shipped\u0026rdquo; notification had already gone out, the compensating action there isn\u0026rsquo;t deletion (you can\u0026rsquo;t unsend an email); it\u0026rsquo;s a second, corrective email.\nThis is exactly why post four\u0026rsquo;s advice to reorder steps to reduce what needs compensating pays off here: award loyalty points only once the order actually dispatches, not right after payment, and the \u0026ldquo;reverse loyalty points\u0026rdquo; compensating transaction never needs to exist at all; that step simply never fired if packaging failed first.\nBecause this is choreography, no single service has a built-in view of \u0026ldquo;what state is order #4521 in right now?\u0026rdquo; Every event in this saga carries a correlation ID, and a dedicated service consumes the full event stream to reconstruct that view, the practical requirement post four flagged as close to essential once you give up a central orchestrator.\nWhy this is worth designing on paper first None of these four decisions were independent. The boundary decisions in step 0 determined who was even allowed to be a participant in the conversation. The communication style in step 1 determined which technology was even sensible in step 2. And the workflow\u0026rsquo;s failure modes in step 3 could only be reasoned about once steps 1 and 2 had already fixed which interactions were synchronous (temporally coupled, need immediate compensating logic) versus event-driven (already loosely coupled, but harder to observe without a correlation ID).\nThe mistake worth avoiding is treating any one of these as a standalone technology choice (\u0026ldquo;we use gRPC\u0026rdquo; or \u0026ldquo;we use Kafka\u0026rdquo;) made once, up front, company-wide. The more durable habit is what we did above: work interaction by interaction, let the boundary and the business requirement dictate the style, and only then pick the technology and the failure-recovery approach that fits what you\u0026rsquo;ve already decided. Applied consistently, that\u0026rsquo;s most of what separates a microservice architecture that stays maintainable from one that quietly turns into a distributed monolith with extra network hops.\n","permalink":"https://hanhpham.vercel.app/posts/microservices-designing-an-order-workflow/","summary":"Boundaries, communication styles, technology, and sagas all sound reasonable in isolation. Here\u0026rsquo;s what it looks like to actually apply all four decisions to one real workflow, in order.","title":"Designing an Order Workflow: Putting the Patterns Together"},{"content":"Most people\u0026rsquo;s mental model of an AI coding tool is still autocomplete with better manners: you type a prompt, it types an answer, you read the answer and decide what to do with it. An agentic assistant is a different shape of tool. Instead of one prompt in, one answer out, it decides its own next step: reads a file to check something, runs a command, looks at the result, changes its plan, tries something else, across as many steps as the task actually needs, without you specifying each one. The unit of work stops being \u0026ldquo;a reply\u0026rdquo; and becomes \u0026ldquo;a completed task.\u0026rdquo;\nThat\u0026rsquo;s the useful part. It\u0026rsquo;s also the part that creates a problem nobody mentions in the pitch: if the same continuous run of reasoning both does the work and decides whether the work is good, you\u0026rsquo;ve built a system with no actual check in it. It can be wrong and confident about being right, for the same underlying reason: it never left its own train of thought.\nWhat \u0026ldquo;agentic\u0026rdquo; looks like in practice Concretely: instead of asking \u0026ldquo;how would I fix this bug?\u0026rdquo; and pasting the answer in yourself, you point an agentic assistant at the actual problem: \u0026ldquo;this fails under X, find out why and fix it\u0026rdquo;, and it goes and reads the relevant code, forms a hypothesis, maybe runs the failing case to confirm it, makes the change, and runs the tests to check. Each of those is a decision it made about what to do next, not a step you told it to take. The value is real: it can hold a much longer chain of \u0026ldquo;check, then act, then check again\u0026rdquo; than pasting snippets back and forth ever allowed.\nThe cost is that all of that now happens inside one continuous context. The hypothesis, the fix, and the \u0026ldquo;yep, looks right\u0026rdquo; are all produced by the same run of reasoning, which means they share the same blind spots by construction. An agent that misunderstood the bug will just as fluently misjudge its own fix as correct; nothing about \u0026ldquo;checking your own work\u0026rdquo; forces you to notice an assumption you didn\u0026rsquo;t know you were making.\nThe fix: don\u0026rsquo;t ask it to grade its own homework The technique that actually addresses this is unglamorous: get a second, independent look, one that wasn\u0026rsquo;t part of producing the answer and isn\u0026rsquo;t primed to agree with it. Concretely, that means starting a fresh session/context (not continuing the one that wrote the fix) and handing it something like this:\nYou are reviewing a proposed fix, not writing one. You were not involved in producing it and don\u0026#39;t know why it was written this way. Original problem: \u0026lt;paste the bug report / requirement / failing test, verbatim\u0026gt; Proposed fix: \u0026lt;paste the diff or final code, verbatim -- not a summary of it\u0026gt; Your job is to find a reason this fix is wrong, incomplete, or solves a different problem than the one described above. Do not confirm it\u0026#39;s correct. Specifically check: - Does it address the root cause, or just the reported symptom? - What input/case would still break, given this fix? - Does it introduce a new problem the original code didn\u0026#39;t have? If you genuinely can\u0026#39;t find a problem after actually trying, say so explicitly and state what you checked -- don\u0026#39;t default to \u0026#34;looks good.\u0026#34; The framing matters as much as the mechanism. \u0026ldquo;Check this is right\u0026rdquo; and \u0026ldquo;try to find what\u0026rsquo;s wrong with this\u0026rdquo; produce noticeably different scrutiny from the same underlying model, on the same input: one invites a skim and a nod, the other invites someone to actually go looking. Handing the second pass only the inputs and the final result, not the trail of reasoning that got there, is what makes it independent rather than a second read of the same argument. Give it the reasoning too, and it tends to inherit the original framing along with it; you get agreement, not verification.\nThis isn\u0026rsquo;t specific to AI; it\u0026rsquo;s the same reason a human code reviewer who only skims for \u0026ldquo;does this look reasonable\u0026rdquo; catches far less than one explicitly asked to find a problem. What\u0026rsquo;s different with an agentic assistant is that the \u0026ldquo;author\u0026rdquo; and the \u0026ldquo;reviewer\u0026rdquo; are trivially the same system unless you deliberately split them, so it\u0026rsquo;s easy to skip this step without noticing you skipped it: there\u0026rsquo;s no separate person you forgot to loop in, just a self-assessment that quietly stood in for one.\nMake it something you don\u0026rsquo;t have to retype Pasting that prompt in by hand every time is how this habit quietly stops happening the moment you\u0026rsquo;re in a hurry. If your assistant supports custom instructions or reusable skills/rules files (most agentic tools do, in one form or another), turn it into one of those instead of a habit you have to remember:\n--- name: adversarial-verify description: Independently check a proposed fix or finding. Use before accepting any non-trivial change as done, especially one you produced yourself in an earlier session. --- ## Rule Never review work in the same context that produced it. Start fresh. ## Inputs to provide - The original problem/requirement, verbatim - The proposed fix/diff/finding, verbatim - Nothing else -- no summary of the reasoning, no \u0026#34;this should be correct\u0026#34; ## What to do Try to find a reason this is wrong or incomplete. Do not confirm it\u0026#39;s correct as your default. Check specifically: does it address the root cause or just the symptom, what case would still break it, and does it introduce a new problem. ## Output State a verdict (confirmed / still broken / unclear) and *why*, not just \u0026#34;looks fine.\u0026#34; An unexamined \u0026#34;looks fine\u0026#34; is not an acceptable output. Having it as a named, reusable thing changes the odds it actually gets used: it turns \u0026ldquo;I should probably double-check this\u0026rdquo; into a single command, which is the difference between a principle and a habit.\nWorking while you\u0026rsquo;re not watching The same thing that makes an agentic assistant able to check its own next step also makes it able to run for a long time without you in the loop at all, hours or overnight. That\u0026rsquo;s a genuinely different mode from interactive back-and-forth, and it fails in different ways if you just hand over a vague goal and walk off.\nWhat actually makes an unattended run trustworthy:\nEvery task is self-contained. No step should assume you\u0026rsquo;re around to clarify. If a brief needs a follow-up question to make sense, that\u0026rsquo;s a sign it wasn\u0026rsquo;t ready to run unattended yet; fix the brief, don\u0026rsquo;t count on being there to answer. Work is split into independent units, not one long dependent chain. A multi-hour run structured as ten separate, self-contained tasks survives one of them going sideways. A single unbroken chain doesn\u0026rsquo;t: one wrong turn early on just compounds for the next nine hours with nobody there to notice. A blocked task says so and stops, instead of guessing. The failure mode to design against isn\u0026rsquo;t \u0026ldquo;it crashed\u0026rdquo;; a crash is at least visible. It\u0026rsquo;s the task that hits genuine ambiguity, picks an interpretation without flagging it, and keeps going as if that were obviously correct. Everything gets written down as it happens, not reconstructed after. You weren\u0026rsquo;t there to see it unfold, so the only thing you have to review in the morning is whatever record it left: a final \u0026ldquo;done\u0026rdquo; with no trail is worthless, since there\u0026rsquo;s no way to tell a real result from a confident guess after the fact. That last point is why the adversarial check from earlier isn\u0026rsquo;t optional for unattended work; it\u0026rsquo;s the whole ballgame. Everything a live session produces, you at least skim as it happens; you develop a rough sense of \u0026ldquo;this seems off\u0026rdquo; in real time even before any formal check. An overnight run gives you none of that. The independent, primed-to-find-problems pass is the only thing standing between \u0026ldquo;it ran all night\u0026rdquo; and \u0026ldquo;it ran all night and produced something real\u0026rdquo;: treat a big pile of unattended output as more in need of adversarial review, not less, exactly because nobody watched it happen.\nWhat this doesn\u0026rsquo;t solve Adversarial verification catches a specific failure mode (confident, self-consistent wrongness), not everything. A second pass that\u0026rsquo;s given a bad spec, or that shares the first pass\u0026rsquo;s actual blind spot rather than just its conclusion, will still miss it. It\u0026rsquo;s a check on the work, not a replacement for understanding the change well enough to know what \u0026ldquo;correct\u0026rdquo; even means here. The decision about whether a fix is actually the right fix, versus merely a fix that survives an adversarial pass, still has to land somewhere, and that somewhere is still you.\nWhat changed for me isn\u0026rsquo;t that I trust agentic tools more now. It\u0026rsquo;s that I stopped treating \u0026ldquo;it said it works\u0026rdquo; as evidence of anything, and started treating a second, independently-framed check as the actual unit of confidence: the same discipline I\u0026rsquo;d want from any collaborator whose reasoning I can\u0026rsquo;t fully see into.\n","permalink":"https://hanhpham.vercel.app/posts/how-ai-actually-fits-into-my-dev-workflow/","summary":"\u0026ldquo;Agentic\u0026rdquo; gets thrown around loosely, but it means something specific: an assistant that decides its own next step instead of waiting for yours. That\u0026rsquo;s also exactly why you can\u0026rsquo;t trust it to tell you when it\u0026rsquo;s done a good job; here\u0026rsquo;s the verification habit that fixes that.","title":"Agentic Coding Assistants, and Why I Don't Let Them Grade Their Own Work"},{"content":"Here\u0026rsquo;s the moment almost every team hits when splitting a monolith: some operation used to be one clean database transaction: say, marking a customer\u0026rsquo;s enrollment VERIFIED and deleting their row from PendingEnrollments, both or neither. Now that logic lives in two separate services with two separate databases. Either change can fail independently of the other, and there\u0026rsquo;s no single ROLLBACK that touches both. The instinct is to reach for some way to make one transaction span both processes. That instinct is usually a mistake, but understanding why tells you what to do instead.\nWhat you actually lose when you split a transaction A normal ACID database transaction gives you four guarantees: atomicity (all changes happen or none do), consistency (the data stays valid), isolation (no one sees an in-progress transaction\u0026rsquo;s intermediate state), and durability (once committed, it\u0026rsquo;s committed).\nYou don\u0026rsquo;t lose all of this when you split a monolith: a single microservice can still use a completely normal ACID transaction for changes to its own database. What you lose is atomicity across microservices. Split the customer/enrollment update into two services with two databases, and you now have two independent transactions, each of which can succeed or fail on its own. There\u0026rsquo;s no wrapper around both that guarantees \u0026ldquo;both or neither.\u0026rdquo;\nThe instinctive fix: two-phase commit, and why it doesn\u0026rsquo;t hold up The obvious next move is a distributed transaction, most commonly implemented as a two-phase commit (2PC). It works in two steps: during the voting phase, a central coordinator asks every participant \u0026ldquo;can you make this change?\u0026rdquo; Each participant locks the relevant resource and promises it can commit later. If everyone votes yes, the coordinator sends a commit message and the changes actually happen; if anyone votes no, everyone rolls back and releases their locks.\nThis sounds reasonable until you look at what it costs:\nDistributed locking, for the duration of the transaction. Every participant holds a lock from the moment it votes yes until it gets the commit message. The longer the transaction, or the more participants involved, the longer those locks are held, and lock contention across multiple services is a much worse problem than lock contention inside one database. A wide window of inconsistency. The coordinator can\u0026rsquo;t guarantee every participant commits at exactly the same instant: the commit message itself has to travel over the network to each one. Isolation, one of the four ACID guarantees, is quietly gone. Nasty failure modes. What happens when a participant votes yes, then goes silent when asked to actually commit? Some of these situations resolve automatically; others need a human to intervene manually. Availability gets worse, not better, as you add participants. Pat Helland\u0026rsquo;s framing of this is the one worth remembering: in most distributed transaction systems, one node failing stalls the whole commit. The more nodes involved, the more likely something is down at any given moment, like an airplane where every additional engine you add is one more way for a required system to fail. 2PC isn\u0026rsquo;t useless in every context: Google\u0026rsquo;s Spanner uses distributed transactional algorithms successfully, but only by controlling the entire stack down to synchronized atomic clocks across data centers, applied within what\u0026rsquo;s logically one database. That\u0026rsquo;s a different problem than coordinating independently-owned microservices, and not a bar most teams need or want to clear.\nFor coordinating state across independently deployed microservices: avoid distributed transactions. If a piece of data genuinely needs true atomic, consistent handling and you can\u0026rsquo;t find a sane way to get that without an ACID transaction, that\u0026rsquo;s a signal the data shouldn\u0026rsquo;t be split apart in the first place: leave it in one service, one database, for now.\nSagas: give up atomicity, gain an explicit process If you do need to coordinate a change across several services (and for anything long-running, you will), the better tool is a saga. The idea predates microservices; it was originally designed for long-lived transactions that would otherwise lock a database for uncomfortably long periods. Instead of one transaction spanning the whole operation, you break it into a sequence of smaller, independent transactions, each with its own local commit.\nThe trade you\u0026rsquo;re making is explicit and important: a saga does not give you atomicity at the level of the whole operation. Each individual step can be a normal, atomic ACID transaction against its own service\u0026rsquo;s database, but there\u0026rsquo;s no wrapper guaranteeing all steps happen together. What a saga gives you instead is enough information to know what state you\u0026rsquo;re in, and it\u0026rsquo;s on you to define what happens next when something goes wrong partway through.\nOne limitation worth internalizing before anything else: a saga handles business failures, not technical failures. \u0026ldquo;The customer\u0026rsquo;s card was declined\u0026rdquo; is a business failure the saga is designed to handle. \u0026ldquo;The payment gateway timed out and threw a 500\u0026rdquo; is a technical failure: the saga assumes the underlying services are fundamentally reliable, and you handle unreliability separately (retries, circuit breakers, that kind of thing), not by asking the saga logic to cover for it.\nRecovering from failure: rolling back vs. pushing forward There are two ways a saga can recover once something goes wrong partway through:\nBackward recovery: undo what\u0026rsquo;s already been committed via compensating transactions, and treat the whole operation as cancelled. Forward recovery: retry the failed step and keep going from where it broke. A single saga can mix both: MusicCorp\u0026rsquo;s order fulfillment might roll the whole order back if an item turns out not to be in the warehouse despite the system thinking it was in stock, but simply retry (queuing for the next day) if the courier has no space today; rolling the entire order back over a shipping delay would be a strange overreaction.\nHere\u0026rsquo;s the part that trips people up the first time: a compensating transaction is not a real rollback. A database rollback happens before commit, and afterward it\u0026rsquo;s as though nothing occurred. A compensating transaction runs after the fact, undoing the effect of something that genuinely did happen. Sometimes that\u0026rsquo;s clean: refund a payment. Sometimes it can\u0026rsquo;t be, because some actions have no true inverse: you cannot unsend an email telling a customer their order shipped. The best you can do is send a second email saying it didn\u0026rsquo;t. This is why these are called semantic rollbacks: they clean up enough for the saga\u0026rsquo;s purposes, without pretending the original action never happened.\nOne low-effort trick that pays for itself: reorder your saga\u0026rsquo;s steps so the parts most likely to fail happen earliest. If awarding loyalty points only happens after the order is actually dispatched, rather than right after payment, you never need a compensating transaction for \u0026ldquo;un-award points\u0026rdquo; in the first place; that step simply never ran if something failed earlier.\nTwo ways to build a saga: orchestration vs. choreography Orchestrated sagas: one conductor An orchestrator (often just a service like OrderProcessor) owns the whole business process, deciding what happens next and calling each downstream service in turn. This tends to lean heavily on request-response: the orchestrator asks Payment to take money, waits, decides what to do based on the result.\nThe upside is visibility: the entire business process is explicitly readable in one place, which is genuinely valuable for onboarding and understanding \u0026ldquo;how does this actually work.\u0026rdquo; The downside is exactly what you\u0026rsquo;d expect from centralizing anything: the orchestrator ends up knowing about, and being coupled to, every service it coordinates, and there\u0026rsquo;s a constant gravitational pull for logic that belongs in those downstream services to instead pile up in the orchestrator. Left unchecked, your services become thin and anemic, and the orchestrator becomes the one place all the actual behavior lives.\nChoreographed sagas: nobody\u0026rsquo;s in charge, trust but verify A choreographed saga distributes the process across the collaborating services themselves, communicating almost entirely through events. When Warehouse sees an Order Placed event, it knows (on its own, with no one telling it to) that its job is to reserve stock and fire an event when that\u0026rsquo;s done. No service needs to know any other service exists; each just reacts to events it cares about.\nThis is a fundamentally more loosely coupled architecture, and it distributes responsibility instead of concentrating it in one place. The cost is that there\u0026rsquo;s no single spot to look at and understand \u0026ldquo;what is the process.\u0026rdquo; You have to reconstruct the whole flow mentally from each service\u0026rsquo;s independent behavior, and you lose an obvious place to ask \u0026ldquo;what state is this specific order\u0026rsquo;s saga in right now?\u0026rdquo;\nThe standard fix for that last problem is a correlation ID: a unique identifier generated for the saga and carried through every event and log line associated with it. With that in place, a dedicated service can consume the event stream and reconstruct a live view of where every saga currently stands; genuinely close to essential if you\u0026rsquo;re going the choreography route.\nWhich one should you pick? The rule of thumb that holds up in practice: orchestration works well when a single team owns the entire process: the extra coupling is easy to manage inside one team\u0026rsquo;s boundary. Choreography earns its extra complexity when multiple teams are involved, because it lets each team own their piece without needing to coordinate through a shared central service. Nothing stops you from mixing styles within one system, or even within a single saga: MusicCorp\u0026rsquo;s fulfillment saga might be choreographed at the top level while Warehouse internally orchestrates its own packaging-and-dispatch sub-flow.\nTwo posts ago we picked communication styles; last post, the technology to implement them; this post, how to keep a multi-step business process consistent when it spans several services. Next, we put all three decisions together and design one real workflow end to end.\n","permalink":"https://hanhpham.vercel.app/posts/microservices-sagas-vs-two-phase-commit/","summary":"Splitting a monolith means splitting its database transactions too. The instinct is to reach for a distributed transaction. Here\u0026rsquo;s why that instinct is usually wrong, and what to do instead.","title":"When Two-Phase Commit Isn't the Answer: Sagas for Microservice Transactions"},{"content":"The first reverse engineering challenge I ever opened was a simple crackme: a tiny binary that prints \u0026ldquo;enter the password\u0026rdquo; and checks your input. The writeup said it was beginner-friendly. I couldn\u0026rsquo;t even find where the password check was. I stared at Ghidra\u0026rsquo;s decompiler output for two hours, saw something like iVar1 = strcmp(local_18, \u0026quot;s3cr3t\u0026quot;) and didn\u0026rsquo;t know what strcmp was, what local_18 was, or why the decompiler showed variable names that looked like a robot named them. I closed the laptop and didn\u0026rsquo;t come back for a week.\nThis is the post I wish someone had written for me. Not a roadmap with 47 resources and a 12-month timeline, just an honest account of what the learning actually looks like, what wastes time, and what actually moves you forward.\nThe beginning is worse than you think Everyone told me to \u0026ldquo;just start doing CTFs.\u0026rdquo; So I opened picoCTF, picked a reversing challenge, and immediately hit a wall. I didn\u0026rsquo;t know how to read assembly. I didn\u0026rsquo;t know what a stack frame was. I didn\u0026rsquo;t know the difference between static and dynamic analysis. I knew how to write Go and Python. That was it.\nThe problem isn\u0026rsquo;t that reverse engineering is impossibly hard. The problem is that there\u0026rsquo;s a prerequisite layer that nobody tells you about (the \u0026ldquo;how does a computer actually work\u0026rdquo; layer), and if you skip it, every tutorial feels like it\u0026rsquo;s assuming knowledge you don\u0026rsquo;t have.\nHere\u0026rsquo;s the prerequisite layer I didn\u0026rsquo;t know I needed:\nWhat a binary actually is. Not \u0026ldquo;a compiled program\u0026rdquo;, but literally what\u0026rsquo;s in the file. ELF headers, sections, the difference between .text and .data and .rodata. You don\u0026rsquo;t need to memorize the ELF spec, but you need to know that readelf -h tells you the entry point and strings gives you the readable text.\nWhat assembly looks like and what it means. Not \u0026ldquo;write assembly\u0026rdquo;; just read it. mov rdi, rax means \u0026ldquo;copy the value in rax into rdi.\u0026rdquo; call means \u0026ldquo;jump to this function.\u0026rdquo; cmp + jne means \u0026ldquo;if these aren\u0026rsquo;t equal, jump somewhere.\u0026rdquo; That\u0026rsquo;s 80% of what you need to start.\nHow function calls work. Arguments go in registers (rdi, rsi, rdx, rcx, r8, r9 on x86-64 Linux). The return value lands in rax. The stack stores local variables and return addresses. If you know this, you can read most Ghidra decompiler output.\nWhat actually helped (and what didn\u0026rsquo;t) Things that helped Compile your own C code and look at the assembly. This was the single biggest breakthrough. Write a ten-line C function (something with an if statement, a loop, a function call), compile it with gcc -O0 -S and read the assembly. Then compile with -O2 and see how it changes. Compiler Explorer (godbolt.org) makes this instant. After doing this fifty times, assembly stops being random characters and starts being readable.\nDo crackmes, not CTFs. CTF reversing challenges are often designed to be tricky: custom VMs, obfuscation, multi-stage protections. Crackmes on crackmes.one are simpler: \u0026ldquo;find the password.\u0026rdquo; That\u0026rsquo;s exactly what you want when you\u0026rsquo;re starting. The goal isn\u0026rsquo;t to be challenged; it\u0026rsquo;s to build pattern recognition. Do thirty easy ones before you touch a CTF.\nUse Ghidra\u0026rsquo;s decompiler, not just the disassembly. Ghidra\u0026rsquo;s decompiler turns assembly into something that looks like C. It\u0026rsquo;s not perfect, but it gives you function names (sometimes), variable types (sometimes), and a readable control flow. Start with the decompiler view. When it\u0026rsquo;s confusing, switch to the disassembly to see what\u0026rsquo;s really happening.\nRead other people\u0026rsquo;s writeups after you solve (or give up on) a challenge. Every writeup teaches you a technique you didn\u0026rsquo;t think of. After reading fifty writeups, you start recognizing patterns: \u0026ldquo;oh, this is a XOR loop,\u0026rdquo; \u0026ldquo;this is checking a hash,\u0026rdquo; \u0026ldquo;this is a custom Base64 encoding.\u0026rdquo; Pattern recognition is 80% of reverse engineering.\nThings that didn\u0026rsquo;t help Trying to learn x86 assembly from a reference manual. The Intel manual is 5,000 pages. You don\u0026rsquo;t need it. You need to know maybe 30 instructions to start: mov, push, pop, call, ret, cmp, je, jne, jmp, xor, add, sub, lea, test, sete, setne. Learn those and you can read most functions.\nSpending a week setting up the \u0026ldquo;perfect\u0026rdquo; lab. FLARE VM, Remnux, Docker containers, custom GDB configs: I spent more time configuring tools than using them. A working Ghidra install and GDB with pwndbg is enough to start. Optimize your setup later.\nTrying to understand every instruction. When you see a function with 200 instructions, you don\u0026rsquo;t need to read all 200. Find the strcmp call. Find the conditional jump after it. Find the \u0026ldquo;success\u0026rdquo; and \u0026ldquo;failure\u0026rdquo; paths. Most of the function is boilerplate you can skip.\nReading about RE without doing RE. Blog posts, YouTube videos, and courses are useful, but only after you\u0026rsquo;ve tried and failed at a challenge. The failure creates the context that makes the learning stick. Read a writeup after you\u0026rsquo;ve spent an hour staring at the binary, not before.\nThe mental model that changed everything The breakthrough for me was realizing that reverse engineering is debugging without source code. When you debug your own code, you set breakpoints, step through execution, watch variables, and form hypotheses about what\u0026rsquo;s wrong. Reverse engineering is the same process; you just don\u0026rsquo;t have variable names and the function names are stripped.\nThe workflow that works:\nRun the binary with junk input. See what happens. What does it print? What files does it create? What network calls does it make? Use strace and ltrace to see syscalls and library calls.\nLook at strings. strings binary | grep -i flag or strings binary | grep -i password. Half the time, the answer is right there. Don\u0026rsquo;t skip the obvious.\nFind the check. In the decompiler, search for strcmp, memcmp, strncmp, or a loop that compares bytes. That\u0026rsquo;s where the validation happens. Follow the string references backwards from there.\nSet a breakpoint at the check. In GDB, break at the comparison. Run the program with a test input. When the breakpoint hits, look at the registers; one register holds your input, the other holds the expected value. There\u0026rsquo;s the password.\nForm a hypothesis, test it. \u0026ldquo;I think this function checks the first character.\u0026rdquo; Set a breakpoint, change the first character, see if the behavior changes. Scientific method, but for binaries.\nThe timeline (honest version) Here\u0026rsquo;s roughly how long each stage took me, working a few hours a week:\nMonth 1: Couldn\u0026rsquo;t read assembly. Didn\u0026rsquo;t know what Ghidra was doing. Solved zero challenges. Read a lot of confusing writeups.\nMonth 2: Could read simple assembly. Could find strcmp in a decompiler. Solved my first three crackmes (all \u0026ldquo;find the string\u0026rdquo; challenges). Felt like I was making progress.\nMonth 3: Could follow control flow. Understood stack frames. Started recognizing common patterns (XOR loops, Base64, simple hashes). Solved easy CTF reversing challenges.\nMonth 4-6: Could tackle medium challenges. Started understanding anti-debugging tricks. Could read C++ binaries (vtables, name mangling). Could write Ghidra scripts to automate repetitive analysis.\nStill working on: Custom VMs, advanced obfuscation, kernel RE, firmware.\nThe gap between \u0026ldquo;I can\u0026rsquo;t read assembly\u0026rdquo; and \u0026ldquo;I can solve easy challenges\u0026rdquo; is smaller than it feels. The gap between \u0026ldquo;I can solve easy challenges\u0026rdquo; and \u0026ldquo;I can tackle anything\u0026rdquo; is enormous and never fully closes. That\u0026rsquo;s what makes it fun.\nResources that actually worked for me crackmes.one: the practice ground. Start at difficulty 1/6. Nightmare (guyinatuxedo.github.io): a course built around CTF challenges, with detailed writeups for each. Ghidra: free, has a real decompiler, and the NSA released it so you know it\u0026rsquo;s not going away. pwndbg: a GDB plugin that makes debugging less painful. Compiler Explorer (godbolt.org): for the compile-and-compare technique that teaches you assembly. CTFtime (ctftime.org): for finding CTFs to practice in. Filter by \u0026ldquo;reverse\u0026rdquo; tag. This is the first post in a series about learning reverse engineering. Next up: your first crackme walkthrough; from downloading the binary to extracting the flag, with every tool click explained.\n","permalink":"https://hanhpham.vercel.app/posts/how-i-started-learning-reverse-engineering/","summary":"I spent my first month with Ghidra staring at assembly I didn\u0026rsquo;t understand, convinced I was too stupid for this. Here\u0026rsquo;s the honest version of how I got from \u0026lsquo;what is a register\u0026rsquo; to solving CTF reversing challenges, and the mistakes that cost me weeks.","title":"How I Started Learning Reverse Engineering (and What I Wish I Knew First)"},{"content":"You attach GDB to a binary. It immediately exits. Or it takes a different path than it did without GDB. Or the flag disappears. The binary knows you\u0026rsquo;re debugging it, and it\u0026rsquo;s fighting back.\nAnti-debugging is a cat-and-mouse game. The binary checks for signs of a debugger, and if it finds one, it takes a different code path, usually one that fails. Your job is to make it think it\u0026rsquo;s not being debugged, or to patch out the checks entirely.\nThis post covers the most common techniques and how to get around them.\nptrace: the classic On Linux, a process can only be debugged by one other process at a time. If the binary calls ptrace(PTRACE_TRACEME) on itself, it\u0026rsquo;s asking to be traced; but if it\u0026rsquo;s already being traced (by GDB), the call fails. If it\u0026rsquo;s NOT being traced, the call succeeds. The binary can check the return value:\nif (ptrace(PTRACE_TRACEME, 0, 0, 0) == -1) { // debugger detected printf(\u0026#34;No debugging allowed!\\n\u0026#34;); exit(1); } Bypass Option 1: Use GDB\u0026rsquo;s catch command to intercept the ptrace call:\n(gdb) catch syscall ptrace (gdb) run # when it hits the catch: (gdb) set $rax = 0 # make ptrace return 0 (success) (gdb) continue Option 2: Patch the binary. Find the ptrace call in Ghidra, NOP it out:\n# in radare2: r2 -w crackme / ptrace # find the call wa nop # replace with nop q Option 3: Use LD_PRELOAD to override ptrace:\n// fake_ptrace.c long ptrace(int request, int pid, void *addr, void *data) { return 0; // always succeed } gcc -shared -o fake_ptrace.so fake_ptrace.c LD_PRELOAD=./fake_ptrace.so ./crackme /proc/self/status The binary reads /proc/self/status and checks the TracerPid line. If a debugger is attached, TracerPid shows the debugger\u0026rsquo;s PID. Otherwise it\u0026rsquo;s 0:\nFILE *f = fopen(\u0026#34;/proc/self/status\u0026#34;, \u0026#34;r\u0026#34;); char line[256]; while (fgets(line, sizeof(line), f)) { if (strncmp(line, \u0026#34;TracerPid:\u0026#34;, 10) == 0) { int pid = atoi(line + 10); if (pid != 0) { // debugger detected } } } Bypass Use a modified kernel or a LD_PRELOAD that intercepts fopen/fgets and returns fake /proc/self/status data where TracerPid: 0.\nOr patch the binary to skip the check: find the comparison and NOP it.\nTiming checks The binary takes a timestamp before and after a section of code. If the difference is too large (because a debugger makes code run slower), it knows it\u0026rsquo;s being debugged:\nstruct timeval start, end; gettimeofday(\u0026amp;start, NULL); // ... some computation ... gettimeofday(\u0026amp;end, NULL); long elapsed = (end.tv_sec - start.tv_sec) * 1000000 + (end.tv_usec - start.tv_usec); if (elapsed \u0026gt; 1000) { // more than 1ms // debugger detected } Bypass Patch the comparison. Or use GDB to set a breakpoint AFTER the timing check, so the timing code runs without interruption.\n/proc/self/maps The binary reads /proc/self/maps to check which libraries are loaded. If it sees GDB\u0026rsquo;s libraries or pwndbg, it knows it\u0026rsquo;s being debugged:\nFILE *f = fopen(\u0026#34;/proc/self/maps\u0026#34;, \u0026#34;r\u0026#34;); char line[256]; while (fgets(line, sizeof(line), f)) { if (strstr(line, \u0026#34;pwndbg\u0026#34;) || strstr(line, \u0026#34;gdb\u0026#34;)) { // debugger detected } } Bypass Patch the string comparison, or use a debugger that doesn\u0026rsquo;t load obvious shared libraries.\nint3 (software breakpoint detection) GDB sets breakpoints by replacing a byte with 0xCC (int3). The binary can scan its own code for 0xCC bytes:\nunsigned char *code = (unsigned char *)main; for (int i = 0; i \u0026lt; 1000; i++) { if (code[i] == 0xCC) { // breakpoint detected } } Bypass Use hardware breakpoints instead of software breakpoints. In GDB:\n(gdb) hbreak check_function # hardware breakpoint Hardware breakpoints use CPU debug registers, not code modification. The binary can\u0026rsquo;t detect them by scanning its own code.\nOr use int3 at a different location: set a breakpoint at the function that checks for breakpoints (metaprogramming, but it works).\nHow to find anti-debugging in a binary When you suspect anti-debugging, look for these strings in the binary:\nstrings crackme | grep -i -E \u0026#34;(debug|ptrace|trace|break|dump|detect)\u0026#34; Common function calls to look for in Ghidra:\nptrace: the classic check fopen with /proc/self/ paths: reading process info gettimeofday or clock_gettime: timing checks IsDebuggerPresent: Windows-specific (PE binaries) CheckRemoteDebuggerPresent: Windows The mindset Anti-debugging isn\u0026rsquo;t magic. It\u0026rsquo;s just more code: code that checks conditions and branches. Every check has a pattern:\nDetect: the binary checks something (ptrace, timing, memory). Decide: if the check fails, take the bad path. Act: crash, exit, or return wrong answer. Your job is to find the detect step and either make it succeed (fake the environment) or remove it (patch the code). Every anti-debugging technique is breakable because the binary runs on your machine and you control the machine.\nThe only question is how much work it takes.\nThis wraps up the introductory RE series. You\u0026rsquo;ve learned to read binaries with strings and Ghidra, solve crackmes, reverse custom VMs, and bypass anti-debugging. The best way to get better is to practice: pick a CTF challenge, set a timer for an hour, and work through it. When you get stuck, look at other people\u0026rsquo;s writeups to learn techniques you missed. The cycle of try, get stuck, read a writeup, try again is how everyone learns this.\n","permalink":"https://hanhpham.vercel.app/posts/anti-debugging-tricks/","summary":"Some binaries detect that you\u0026rsquo;re debugging them and crash, change behavior, or hide the flag. This post covers the common anti-debugging techniques and how to bypass them.","title":"Anti-Debugging Tricks: When the Binary Fights Back"},{"content":"Some CTF reversing challenges don\u0026rsquo;t check a password. Instead, they load your input into a custom virtual machine and run bytecode. The bytecode is designed to be hard to follow: obfuscated instruction sets, indirect jumps, data manipulation that makes no sense until you understand the VM\u0026rsquo;s design. But once you reverse the instruction set, you can write a solver that runs backwards or brute-forces the flag.\nThis post covers how to spot a custom VM, how to reverse its instruction set, and how to write a solver.\nWhat is a custom VM? A virtual machine in this context is a program that:\nReads bytecode (an array of bytes) from somewhere: the binary itself, a file, or hardcoded in the data section. Interprets each byte as an instruction using a switch/case or jump table. Executes the instruction using registers, a stack, or memory that only the VM knows about. It\u0026rsquo;s the same idea as Python or Java\u0026rsquo;s VM, but the instruction set is invented for this one challenge. The flag is usually encoded by the VM executing a specific program; your input must make the VM reach a \u0026ldquo;success\u0026rdquo; state.\nHow to spot a VM When you open a binary in Ghidra and see this pattern, it\u0026rsquo;s probably a VM:\nA big switch statement in main (or a called function):\nwhile (pc \u0026lt; code_length) { switch (bytecode[pc]) { case 0x01: /* instruction 1 */ break; case 0x02: /* instruction 2 */ break; case 0x03: /* instruction 3 */ break; // ... 20-50 cases } pc++; } The pc is a program counter: an index into the bytecode array. The switch dispatches on the opcode (first byte of each instruction). The number of cases tells you how many instructions the VM has.\nOther telltale signs:\nA while or for loop that indexes into a byte array and dispatches on the value. Registers that aren\u0026rsquo;t CPU registers: global variables used as VM state (vm_reg[0], vm_stack, vm_pc). A function that looks like it has no purpose other than \u0026ldquo;run this program.\u0026rdquo; Reversing the instruction set Once you\u0026rsquo;ve found the VM loop, the job is to figure out what each opcode does. Work through the switch cases one at a time.\nStart with the obvious ones Some opcodes are easy to identify:\nStack push:\ncase 0x10: vm_stack[vm_sp] = vm_reg[arg]; vm_sp++; break; This pushes a register value onto the stack. You can tell from the vm_sp++ (stack pointer increment) and the write to the stack array.\nStack pop:\ncase 0x11: vm_sp--; vm_reg[arg] = vm_stack[vm_sp]; break; The reverse: decrement the stack pointer, read into a register.\nInput:\ncase 0x20: vm_reg[0] = user_input[input_index]; input_index++; break; Reads a byte from user input into a register. This is how your input gets into the VM.\nCompare and jump:\ncase 0x30: if (vm_reg[arg1] != vm_reg[arg2]) { vm_pc = jump_target; } break; Conditional jump: if two registers don\u0026rsquo;t match, jump to a different offset. This is the \u0026ldquo;check\u0026rdquo; opcode.\nPrint success:\ncase 0xFF: puts(\u0026#34;Correct!\u0026#34;); break; The exit opcode.\nMap the full instruction set Create a table (even just a text file) mapping each opcode to what it does:\n0x00 = nop 0x01 = push reg 0x02 = pop reg 0x03 = mov reg, imm 0x04 = add reg, reg 0x05 = sub reg, reg 0x06 = xor reg, reg 0x10 = read_input reg 0x20 = cmp reg, reg 0x21 = je offset 0x22 = jne offset 0xFF = halt The exact opcodes depend on the challenge. But the categories are consistent: data movement, arithmetic, logic, control flow, I/O. Every VM implements some subset of these.\nFollow the data flow Once you know the instructions, trace what happens to user input:\nIt enters via the read_input opcode. It gets pushed onto the stack or moved into a register. It gets transformed: XOR, ADD, SUB, lookup table, bit shifts. The result is compared against something: the expected flag bytes. The transformation chain is the \u0026ldquo;encryption\u0026rdquo; on the flag. If you can read each step, you can reverse it.\nWriting a solver Once you know the instruction set, you don\u0026rsquo;t need to trace through the VM manually. Write a solver.\nApproach 1: emulate the VM Reimplement the VM in Python (or any language). Run the bytecode through your emulator with candidate inputs:\nclass VM: def __init__(self, bytecode): self.bytecode = bytecode self.pc = 0 self.reg = [0] * 8 self.stack = [] self.input_idx = 0 def run(self, user_input): self.user_input = user_input while self.pc \u0026lt; len(self.bytecode): op = self.bytecode[self.pc] self.pc += 1 if op == 0x10: # read_input arg = self.bytecode[self.pc]; self.pc += 1 self.reg[arg] = self.user_input[self.input_idx] self.input_idx += 1 elif op == 0x06: # xor reg, reg a = self.bytecode[self.pc]; self.pc += 1 b = self.bytecode[self.pc]; self.pc += 1 self.reg[a] ^= self.reg[b] elif op == 0x20: # cmp a = self.bytecode[self.pc]; self.pc += 1 b = self.bytecode[self.pc]; self.pc += 1 if self.reg[a] != self.reg[b]: self.pc = self.bytecode[self.pc] else: self.pc += 1 elif op == 0xFF: return True return False Then brute-force or constraint-solve the input.\nApproach 2: reverse the operations If the VM does simple transformations, reverse them step by step:\n# VM does: input[i] ^ 0x42 + 3 == expected[i] # Reverse: flag[i] = (expected[i] - 3) ^ 0x42 flag = bytes([(b - 3) ^ 0x42 for b in expected]) This only works when the transformations are simple and reversible. If the VM does a hash or a complex loop, you need the emulator approach.\nApproach 3: use angr angr is a binary analysis framework that can solve constraining problems. If you can identify the VM\u0026rsquo;s check function, angr can find an input that makes it return \u0026ldquo;success\u0026rdquo;:\nimport angr proj = angr.Project(\u0026#39;./crackme\u0026#39;, auto_load_libs=False) state = proj.factory.entry_state() simgr = proj.factory.simulation_manager(state) # find the \u0026#34;correct\u0026#34; state, avoid the \u0026#34;wrong\u0026#34; state simgr.explore(find=0x401234, avoid=0x401256) if simgr.found: found = simgr.found[0] print(found.posix.dumps(0)) # stdin that reaches the find address angr treats the binary as a set of constraints and uses a constraint solver to find an input that reaches a specific address. It\u0026rsquo;s slow but it works, even on VMs you don\u0026rsquo;t fully understand.\nCommon VM patterns in CTF Stack-based VMs look like Forth or JVM bytecode. Instructions push and pop values, operations work on the top of the stack:\npush 0x41 push 0x42 xor // stack = [0x03] Register-based VMs look like assembly. Instructions move data between registers and perform operations:\nmov r0, 0x41 mov r1, 0x42 xor r0, r1 // r0 = 0x03 Obfuscated VMs hide the instruction set. Instead of a clean switch, they use computed jumps (goto *(\u0026amp;jump_table + opcode * 8)), so the decompiler can\u0026rsquo;t show you the switch directly. You need to follow the jump table manually. In Ghidra, look for arrays of function pointers in the data section.\nMulti-stage VMs run one VM, then use its output as input to another VM. You need to reverse both stages.\nCustom VMs look intimidating but they follow patterns. Every VM has a program counter, an instruction stream, and a dispatch mechanism. Find those three things and you can reverse the instruction set. The next post covers anti-debugging: techniques binaries use to detect that you\u0026rsquo;re debugging them, and how to get around it.\n","permalink":"https://hanhpham.vercel.app/posts/custom-vms-in-ctf/","summary":"Some CTF reversing challenges hide the flag behind a custom virtual machine: a program that interprets its own bytecode. This post shows you how to spot them, reverse the instruction set, and write a solver.","title":"Custom VMs in CTF: When the Binary Speaks Its Own Language"},{"content":"If you\u0026rsquo;ve never reversed a binary before, this post walks through the entire process on a simple crackme: a small program that asks for a password and tells you if you got it right. Every command, every tool, and every decision is explained. By the end, you\u0026rsquo;ll have a workflow you can repeat on harder challenges.\nThis isn\u0026rsquo;t a walkthrough of a specific real crackme; it\u0026rsquo;s a composite of the patterns you\u0026rsquo;ll see in dozens of beginner challenges. The techniques apply to any simple \u0026ldquo;find the password\u0026rdquo; binary.\nStep 0: download and run it Download the binary from crackmes.one (or wherever your challenge comes from). First thing: what is it?\nfile crackme crackme: ELF 64-bit LSB executable, x86-64, version 1 (SYSV), dynamically linked, interpreter /lib64/ld-linux-x86-64.so.2, BuildID[sha1]=..., for GNU/Linux 3.2.0, not stripped It\u0026rsquo;s a 64-bit Linux executable. \u0026ldquo;Not stripped\u0026rdquo; means function names are still in the binary, which makes things easier. \u0026ldquo;Dynamically linked\u0026rdquo; means it uses shared libraries (libc).\nRun it with junk input:\n./crackme Enter the password: hello Wrong password. Try again. Simple. It asks for a password, you type something, it says wrong. Our job is to find the right one.\nStep 1: strings, the 30-second check Before opening any tool, run strings:\nstrings crackme | head -50 Look for anything interesting: flag formats (flag{, CTF{), readable passwords, error messages, function names. In many beginner crackmes, the password is literally sitting in the strings:\nstrings crackme | grep -i password Enter the password: Wrong password. Try again. Congratulations! You got it right. s3cr3t_p4ss If you see a string that looks like a password (s3cr3t_p4ss), try it:\nEnter the password: s3cr3t_p4ss Congratulations! You got it right. Done. Half the beginner crackmes on crackmes.one are solvable with strings. Don\u0026rsquo;t skip this step; it\u0026rsquo;s embarrassing how often it works.\nIf strings doesn\u0026rsquo;t give you the answer, keep reading.\nStep 2: find the check in Ghidra Open the binary in Ghidra. Create a new project, import the binary, and let it analyze (say yes to all the analysis options). The decompiler view will open automatically.\nFinding the main function Ghidra\u0026rsquo;s symbol tree shows all function names. Look for main; that\u0026rsquo;s where the program starts. Double-click it and the decompiler shows something like:\nundefined8 main(void) { char local_18 [24]; puts(\u0026#34;Enter the password:\u0026#34;); fgets(local_18, 20, stdin); check_password(local_18); return 0; } The password goes into local_18, then check_password is called. That\u0026rsquo;s where the comparison happens.\nFollowing the check Double-click check_password in the decompiler:\nvoid check_password(char *input) { int iVar1; iVar1 = strcmp(input, \u0026#34;d4nger_d4t4\u0026#34;); if (iVar1 == 0) { puts(\u0026#34;Congratulations! You got it right.\u0026#34;); } else { puts(\u0026#34;Wrong password. Try again.\u0026#34;); } } There it is. strcmp compares your input against \u0026quot;d4nger_d4t4\u0026quot;. If they\u0026rsquo;re equal (iVar1 == 0), you win.\nTry it:\nEnter the password: d4nger_d4t4 Congratulations! You got it right. What if it\u0026rsquo;s not a plain strcmp? Beginner crackmes sometimes make it slightly harder: instead of comparing against a string directly, they compare against an XOR-encoded version, or build the string character by character. The pattern to look for in the decompiler:\nXOR loop: a function that XORs each byte of your input against a key and compares the result. In the decompiler, it looks like a loop with ^ (XOR):\nvoid check_password(char *input) { char expected[] = {0x1c, 0x0a, 0x12, 0x06, 0x3a, 0x00}; for (int i = 0; i \u0026lt; 6; i++) { if ((input[i] ^ 0x42) != expected[i]) { puts(\u0026#34;Wrong password.\u0026#34;); return; } } puts(\u0026#34;Correct!\u0026#34;); } To solve: XOR each expected byte against 0x42:\nexpected = [0x1c, 0x0a, 0x12, 0x06, 0x3a, 0x00] key = 0x42 password = \u0026#34;\u0026#34;.join(chr(b ^ key) for b in expected) print(password) # \u0026#34;h3ll0\u0026#34; Character-by-character comparison: the decompiler shows individual comparisons for each character:\nif (input[0] != \u0026#39;f\u0026#39;) { puts(\u0026#34;Wrong.\u0026#34;); return; } if (input[1] != \u0026#39;l\u0026#39;) { puts(\u0026#34;Wrong.\u0026#34;); return; } if (input[2] != \u0026#39;a\u0026#39;) { puts(\u0026#34;Wrong.\u0026#34;); return; } // ... Just read the characters off: flag{...}.\nStep 3: see it happening in GDB Ghidra tells you what the program does. GDB lets you watch it do it. This is where reverse engineering becomes debugging without source code.\nSetup Install pwndbg (a GDB plugin that makes GDB bearable):\npip install pwndbg Run GDB with the binary:\ngdb ./crackme Break at the check If you found the check function name in Ghidra (like check_password), break there:\n(gdb) break check_password If the function is stripped (no name), break at strcmp instead:\n(gdb) break strcmp Run and inspect (gdb) run The program starts, asks for a password, you type test. When it hits the breakpoint:\nBreakpoint 1, check_password (input=0x7fffffffe040 \u0026#34;test\\n\u0026#34;) pwndbg automatically shows you the registers and the arguments. input points to your input string. One of the other registers or stack values holds the expected password.\nLook at the arguments to strcmp:\n(gdb) x/s $rdi 0x7fffffffe040: \u0026#34;test\\n\u0026#34; (gdb) x/s $rsi 0x402000: \u0026#34;d4nger_d4t4\u0026#34; $rdi is your input. $rsi is the expected value. There\u0026rsquo;s the password without even reading the decompiler.\nPatching (advanced) If you want to see what happens on success without knowing the password, you can patch the binary. In GDB:\n(gdb) break check_password (gdb) run # when it hits the breakpoint: (gdb) set $rdi = 0 (gdb) continue This forces strcmp to return 0 (equal), so the program always takes the success path. The binary on disk isn\u0026rsquo;t changed, only the running instance.\nTo actually patch the binary file, use a hex editor or radare2:\nr2 -w crackme # in radare2: wa nop # at the conditional jump after strcmp q Step 4: the flag In CTF challenges, the password often IS the flag, or the flag is printed after you enter the correct password:\nEnter the password: d4nger_d4t4 Congratulations! You got it right. Flag: CTF{y0u_f0und_th3_p4ssw0rd} Sometimes the flag is built dynamically, assembled from pieces stored in different parts of the binary. In that case, you\u0026rsquo;ll see the pieces in Ghidra and reconstruct them:\n# from the decompiler: flag = \u0026#34;CTF{\u0026#34; + chr(0x79) + chr(0x30) + chr(0x75) + \u0026#34;_f0und_th3_p4ssw0rd}\u0026#34; print(flag) # CTF{y0u_f0und_th3_p4ssw0rd} The checklist Every time you open a new crackme, follow these steps in order:\nfile: what kind of binary is this? Run with junk input: what does it do? What does it print? strings: is the answer right there? Ghidra: find the comparison. Is it strcmp? XOR loop? Hash check? GDB: break at the comparison, inspect the registers, read the expected value. Solve: extract the password, submit the flag. If step 3 works, you\u0026rsquo;re done in two minutes. If not, steps 4-5 will get you there. Steps 1-2 are always worth doing first; they tell you what you\u0026rsquo;re dealing with.\nThis workflow works for every beginner crackme and most intermediate ones. As challenges get harder, you\u0026rsquo;ll add techniques (anti-debugging bypasses, custom VM emulation, obfuscation), but the core process stays the same: run it, find the check, read the comparison, extract the answer. The next post covers Ghidra in depth: how to read its decompiler output confidently and use it to understand binaries you\u0026rsquo;ve never seen before.\n","permalink":"https://hanhpham.vercel.app/posts/your-first-crackme-walkthrough/","summary":"A crackme is a small binary that asks for a password and checks your input. This post walks through solving one, from downloading the binary to extracting the flag, with every tool click explained so you can follow along.","title":"Your First Crackme: A Walkthrough from strings to Flag"},{"content":"Last time we settled on the style of communication before touching any technology: sync vs. async, request vs. broadcast. That was the point: pick \u0026ldquo;gRPC\u0026rdquo; or \u0026ldquo;Kafka\u0026rdquo; first and you\u0026rsquo;ve quietly locked in a style whether you meant to or not. With the style decided, the technology conversation gets a lot shorter. This post walks through what MusicCorp would actually reach for, and where each option\u0026rsquo;s real trade-offs sit.\nBefore the specific technologies, five things worth wanting from any choice you make here: make backward-compatible changes easy, make the interface explicit (schemas help), keep the API technology-agnostic so you aren\u0026rsquo;t locking every consumer into your language, make the service cheap for consumers to use, and don\u0026rsquo;t leak internal implementation detail through the wire format. Keep these in the back of your mind as we go through the options; they\u0026rsquo;re the yardstick.\nRemote Procedure Calls: gRPC and the ghosts of RMI RPC\u0026rsquo;s whole pitch is making a network call look like a local method call. That\u0026rsquo;s also its biggest risk: hide the network too well, and developers write code that fires off a thousand \u0026ldquo;local-looking\u0026rdquo; calls without realizing each one is a network round trip.\nOlder RPC implementations like Java RMI compound this with genuine brittleness. Tie your client and server to the same binary stub generation, and removing so much as an unused field from a shared type can break deserialization on every consumer: you\u0026rsquo;re stuck doing lockstep releases whether you wanted to or not.\ngRPC is the modern answer and it\u0026rsquo;s a good one. Built on HTTP/2 with protocol buffers for serialization, it has strong cross-language support (so you don\u0026rsquo;t inherit RMI\u0026rsquo;s single-platform trap), solid performance, and a healthy tooling ecosystem for schema evolution. If you have good control over both ends of a synchronous request-response call (which is gRPC\u0026rsquo;s sweet spot), it\u0026rsquo;s usually the first thing worth evaluating.\nREST: the sensible default, HATEOAS aside REST\u0026rsquo;s actual contribution isn\u0026rsquo;t \u0026ldquo;use HTTP,\u0026rdquo; it\u0026rsquo;s the idea of resources (a Customer, an Order) with a uniform set of verbs (GET, POST, PUT, DELETE) that behave consistently across every resource, instead of a bespoke createCustomer/editCustomer method per operation. Riding on HTTP gets you a huge amount for free: caching proxies, load balancers, monitoring tooling, and a security ecosystem that already understands the protocol.\nThe purist version of REST also includes HATEOAS: hypermedia controls that let a client navigate an API the way a human navigates a website, without hardcoding URLs. It\u0026rsquo;s a genuinely interesting idea. It\u0026rsquo;s also one that, across the industry, essentially never gets adopted in practice; you\u0026rsquo;re unlikely to meet a team that\u0026rsquo;s found it worth the extra plumbing. Don\u0026rsquo;t feel behind if you skip it.\nThe realistic trade-offs: REST-over-HTTP payloads are heavier than a lean binary format, and TCP-based HTTP has more overhead than protocols built to skip it. None of that matters for the overwhelming majority of service-to-service traffic. REST over HTTP remains the sensible default whenever you want maximum interoperability and don\u0026rsquo;t have a specific reason to reach for something else.\nGraphQL: built for one job, not a general replacement GraphQL solves a specific, narrow problem extremely well: letting a constrained client (think a mobile app) issue one query that pulls back exactly the fields it needs from potentially multiple sources, instead of several round trips returning more than it asked for.\nThat\u0026rsquo;s a perimeter-facing job, not a microservice-to-microservice one. Two practical limitations worth knowing up front: caching is much harder than with plain REST, since you can\u0026rsquo;t just slap standard HTTP cache headers on an arbitrary query; and writes don\u0026rsquo;t fit the model nearly as naturally as reads do, which is why teams commonly end up using GraphQL for reads and REST for writes on the same system. If you\u0026rsquo;re aggregating and filtering data for a UI, look at GraphQL or the Backend-for-Frontend pattern, not at replacing your internal service-to-service protocol with it.\nMessage brokers: what \u0026ldquo;guaranteed delivery\u0026rdquo; actually buys you For asynchronous communication, brokers (RabbitMQ, ActiveMQ, Kafka, or a managed equivalent like SQS/SNS) sit in the middle so producers and consumers never have to be up at the same moment. The feature that matters most is guaranteed delivery: the broker holds a message durably until it can be delivered, so the sender doesn\u0026rsquo;t have to decide \u0026ldquo;retry or give up?\u0026rdquo; the way it would with a direct synchronous call.\nTwo structural concepts worth keeping straight:\nQueues are point-to-point: one message, consumed by one member of a consumer group. This is your load-distribution mechanism (the competing consumers pattern): three instances of OrderProcessor in the same group, and only one of them handles any given message. Topics let multiple, independent consumer groups each get their own copy of the same message. This is what event broadcast actually runs on: Warehouse and Notifications both react to the same Order Placed event without knowing about each other. As a rough (not absolute) rule: topics fit event-driven collaboration, queues fit request-response.\nKafka deserves a specific mention because of two features that set it apart from a \u0026ldquo;normal\u0026rdquo; broker. First, message permanence: Kafka can retain messages far longer than \u0026ldquo;until the last consumer reads it,\u0026rdquo; which means a newly deployed consumer can replay history it never saw the first time around. Second, built-in stream processing (KSQL), letting you define SQL-like queries over topics directly, which starts to look like a continuously updating materialized view with a topic as the source instead of a table.\nWatch out for \u0026ldquo;exactly-once delivery\u0026rdquo; claims. It\u0026rsquo;s a genuinely contested topic even among distributed-systems experts: some say it\u0026rsquo;s provably impossible in the general case, others say a few specific workarounds get you there. Whatever your broker claims, build consumers that are idempotent and can tolerate seeing the same message twice (a message ID and a \u0026ldquo;have I processed this already?\u0026rdquo; check goes a long way), rather than betting your correctness on the broker\u0026rsquo;s marketing copy.\nFinding services: DNS, or something built for constant churn Once you have more than a handful of services, you need a way to answer \u0026ldquo;where is Accounts right now?\u0026rdquo; DNS is the simplest starting point (well understood, supported everywhere), but it\u0026rsquo;s designed for a world where hosts don\u0026rsquo;t change every few minutes. Time-to-live caching means clients can hold stale entries, and DNS itself has no good story for \u0026ldquo;this instance just died, stop routing to it\u0026rdquo; without pointing entries at a load balancer that handles that churn for you.\nFor environments where instances come and go constantly, dynamic service registries (Consul, or whatever your orchestration platform provides; Kubernetes ships with its own service discovery via etcd) handle registration and health checking directly, and are generally the better fit once you\u0026rsquo;re past a handful of long-lived instances.\nAPI gateways and service meshes are not the same thing This pairing causes more confused architecture decisions than almost anything else in the space, so it\u0026rsquo;s worth being precise: an API gateway manages north-south traffic, requests entering your system from the outside world. A service mesh manages east-west traffic: service-to-service calls inside your perimeter. They can overlap in practice, but they solve different problems.\nThe API gateway\u0026rsquo;s job, in the overwhelming majority of real systems, is much narrower than the vendor marketing around \u0026ldquo;the API economy\u0026rdquo; suggests: mostly it\u0026rsquo;s mapping external requests (from your own web/mobile clients) to internal services, handling API keys, rate limiting, and logging at the edge.\nThe two misuses I\u0026rsquo;d actively steer you away from: using the gateway for call aggregation (that\u0026rsquo;s a job for GraphQL or a Backend-for-Frontend, not a proxy layer), and using it for protocol rewriting (\u0026ldquo;turn any SOAP API into REST!\u0026rdquo;). Both push business logic into a third-party tool that was never built to hold it; keep the pipes dumb, keep the smarts in your own code.\nA quick note on sharing code DRY is good advice inside a service. Across service boundaries, it needs qualifying: shared libraries are fine for things invisible to the outside world (a logging library), risky the moment they leak into the wire contract. If two services share a library that defines the shape of data sent over the network, you\u0026rsquo;ve reintroduced coupling through the back door: a schema change now means redeploying every consumer of that library, which is exactly the lockstep-release problem microservices are supposed to avoid. If you want a client library at all, keep the consumer in control of when they upgrade it, the way public SDKs (AWS\u0026rsquo;s, for instance) do it. And for the broader question of how to evolve APIs without breaking running consumers (the expand-contract pattern, schema evolution with protobuf, and consumer-driven contracts), see versioning without breaking everyone.\nWe\u0026rsquo;ve now covered how services talk and what to build that communication on. But talking to each other is only half the workflow problem: the harder question is what happens to consistency when a single business operation spans several of these calls and one of them fails partway through. That\u0026rsquo;s next.\n","permalink":"https://hanhpham.vercel.app/posts/microservices-picking-communication-tech/","summary":"The communication style (sync/async, request-response/event-driven) narrows the field. Here\u0026rsquo;s how the actual technologies stack up once you\u0026rsquo;re choosing among them.","title":"REST, gRPC, GraphQL, or a Broker: Picking Your Communication Tech"},{"content":"Ghidra is the free decompiler from the NSA. You open a binary, it shows you C-like code. The problem: the code looks like it was written by someone who hates you. Variables are named iVar1 and local_48. Functions are named FUN_00401230. There are casts everywhere. Nothing has a type. You know it\u0026rsquo;s showing you what the program does, but you can\u0026rsquo;t read it.\nThis post teaches you to read Ghidra\u0026rsquo;s output in 30 minutes. Not every feature, just the 20% you\u0026rsquo;ll use 80% of the time.\nWhy Ghidra looks weird Ghidra\u0026rsquo;s decompiler doesn\u0026rsquo;t have source code. It\u0026rsquo;s reconstructing C from machine code. Machine code doesn\u0026rsquo;t have variable names, types, or comments. So Ghidra makes up names based on what it can figure out:\niVar1: an integer variable. The i prefix means integer, v is the variable, 1 is the number Ghidra assigned. It does this because it doesn\u0026rsquo;t know what you named this variable. local_48: a variable on the stack at offset 0x48 from the stack pointer. It\u0026rsquo;s a local variable, and Ghidra only knows its position, not its purpose. FUN_00401230: a function at address 0x00401230. No symbol name available, so Ghidra uses the address. UNWARNING or WARNING: Ghidra found something suspicious in the decompilation. Usually a cast or a type mismatch. None of these names are permanent. You rename them. That\u0026rsquo;s the first thing to learn.\nThe five things to do immediately 1. Rename variables Click a variable name in the decompiler view, press L, type a new name. If you figure out that local_48 is a buffer holding user input, rename it user_input. If iVar1 is the result of a strcmp, rename it strcmp_result.\nThis is the single most important Ghidra skill. The decompiler output becomes readable the moment you rename things. Do it constantly: every time you understand what something is, name it.\n2. Set types Right-click a variable, choose \u0026ldquo;Retype Variable.\u0026rdquo; If Ghidra thinks local_48 is char[24] but you know it\u0026rsquo;s an int, change it. If a function parameter is shown as void * but you know it\u0026rsquo;s a FILE *, fix it.\nCorrect types make the decompiler output dramatically more readable. When Ghidra knows something is a struct, it shows field access as ptr-\u0026gt;field_name instead of *(ptr + offset). That\u0026rsquo;s a huge difference.\n3. Use the string references Press Shift+F12 to open the strings window. Find an interesting string: an error message, a format string, a URL. Double-click it. Press Ctrl+X to see cross-references: every function that uses this string. This is the fastest way to find the interesting parts of a binary.\nStrings are your map. \u0026ldquo;Wrong password\u0026rdquo; leads to the check function. \u0026ldquo;Flag:\u0026rdquo; leads to the flag output. \u0026ldquo;Usage: %s\u0026rdquo; leads to the argument parser.\n4. Use the function graph Select a function, press Space to switch to the graph view. This shows the control flow as a visual flowchart: diamonds for conditionals, boxes for code blocks, arrows for jumps.\nIf the decompiler output is confusing, the graph view often makes the structure clear. You can see the if/else branches, the loop boundaries, and the success/failure paths as a picture instead of text.\n5. Add comments Select a line, press / to add a comment. When you figure out that a section of code is \u0026ldquo;decoding the flag,\u0026rdquo; comment it. When you realize a function is \u0026ldquo;custom Base64 decoder,\u0026rdquo; comment it. You will forget what you learned if you don\u0026rsquo;t write it down. Ghidra projects save your annotations.\nReading decompiler output: patterns to recognize The strcmp pattern iVar1 = strcmp(input, \u0026#34;expected_password\u0026#34;); if (iVar1 != 0) { puts(\u0026#34;Wrong password\u0026#34;); // failure path } else { puts(\u0026#34;Correct!\u0026#34;); // success path } This is the simplest check. Your input goes into the first argument, the expected value is the second. If you see strcmp, you\u0026rsquo;re done; the second argument is the password.\nThe XOR decode pattern void decode(char *buf, int len) { for (int i = 0; i \u0026lt; len; i++) { buf[i] = buf[i] ^ 0x37; } } A loop that XORs each byte. The 0x37 is the key. To decode: XOR each byte of the encoded data against 0x37. In Python:\nencoded = bytes([0x50, 0x54, 0x57, 0x56, 0x51]) key = 0x37 print(bytes([b ^ key for b in encoded])) # b\u0026#34;hello\u0026#34; The character array pattern local_28[0] = 0x66; local_28[1] = 0x6c; local_28[2] = 0x61; local_28[3] = 0x67; local_28[4] = 0x7b; The flag is being built one byte at a time. Read the hex values as ASCII: 0x66 = f, 0x6c = l, 0x61 = a, 0x67 = g, 0x7b = {.\nThe loop with a lookup table char table[] = \u0026#34;abcdefghijklmnopqrstuvwxyz\u0026#34;; for (int i = 0; i \u0026lt; len; i++) { result[i] = table[input[i] - \u0026#39;a\u0026#39;]; } This is a substitution cipher: each input character is mapped through a table. If the table is the normal alphabet, it\u0026rsquo;s doing nothing. If the table is shuffled, it\u0026rsquo;s a simple substitution.\nThe hash-and-compare pattern uint32_t h = 0x811c9dc5; for (int i = 0; i \u0026lt; len; i++) { h ^= input[i]; h *= 0x01000193; } if (h == 0xdeadbeef) { puts(\u0026#34;Correct!\u0026#34;); } This is a hash function (FNV-1a in this case). The input is hashed, and the hash is compared against a hardcoded value. You can\u0026rsquo;t reverse a hash, but you can use the decompiler to find the hash algorithm, then brute-force or dictionary-attack it.\nThe Ghidra workflow checklist Open the binary, say yes to analysis. Let Ghidra do its thing. Look at the strings (Shift+F12). Find the interesting ones. Follow cross-references from strings to functions. This leads you to the important code. Rename the function based on what it does (check_flag, decode, print_flag). Rename variables inside the function. Every time you figure out what something is, name it. Set types where Ghidra guessed wrong. Add comments for things you\u0026rsquo;ll forget. Read the decompiler output with your renamed, typed, commented code. It\u0026rsquo;s now readable. Do this for ten functions and Ghidra goes from \u0026ldquo;confusing mess\u0026rdquo; to \u0026ldquo;actually useful tool.\u0026rdquo; The investment in renaming pays for itself immediately; you\u0026rsquo;ll never go back to reading unnamed iVar1 variables.\nCommon gotchas Ghidra\u0026rsquo;s decompiler isn\u0026rsquo;t always right. It reconstructs types from machine code, and sometimes it guesses wrong. If something looks weird (a function taking 12 parameters, a nonsensical cast, a loop that doesn\u0026rsquo;t make sense), check the disassembly view. The decompiler is a convenience layer; the disassembly is what the CPU actually sees.\nOptimized code looks different. If the binary was compiled with -O2 or -O3, the compiler inlined functions, reordered instructions, and eliminated \u0026ldquo;dead\u0026rdquo; code. The decompiler output is harder to read because the compiler did things you wouldn\u0026rsquo;t expect. Try -O0 binaries first.\nC++ binaries are harder. Name mangling (_ZN3Foo3barEv), vtables (function pointers in structs), and template instantiation make C++ binaries more complex. Ghidra can demangle names and identify vtables, but you need to understand how C++ compiles to really read them.\nStripped binaries have no names. \u0026ldquo;Stripped\u0026rdquo; means the function names were removed. Ghidra shows FUN_00401230 instead of main. The code is still there; you just have to figure out which function is which by looking at what they do. Entry points (like main) are still findable through the ELF header.\nGhidra is the tool that turns \u0026ldquo;I can\u0026rsquo;t read assembly\u0026rdquo; into \u0026ldquo;I can read this.\u0026rdquo; The first time you rename a variable and the decompiler output suddenly makes sense, you\u0026rsquo;ll understand why people use it. The next post covers a more advanced topic: the custom virtual machines that show up in CTF reversing challenges, and how to reverse them.\n","permalink":"https://hanhpham.vercel.app/posts/ghidra-in-30-minutes/","summary":"Ghidra\u0026rsquo;s decompiler turns machine code into something that looks like C. But the output is full of strange variable names, weird casts, and functions that look nothing like what you\u0026rsquo;d write. Here\u0026rsquo;s how to read it anyway.","title":"Ghidra in 30 Minutes: Reading Assembly Without Hating It"},{"content":"You\u0026rsquo;ve read the series on boundaries, communication, and sagas. Your services are well-bounded, your communication patterns are deliberate, and your saga handles the order workflow cleanly. Then Monday morning, a customer reports that their order disappeared. The payment went through (the payment service confirms it), but the warehouse never got the stock reservation request. Where did it go?\nIn a monolith, the answer is in one stack trace. In microservices, that stack trace doesn\u0026rsquo;t exist. The request lived across three services, each with its own logs, each running in its own process. You now have a debugging problem that\u0026rsquo;s fundamentally different from anything a monolith teaches you, and if you don\u0026rsquo;t prepare for it in advance, your first real incident will be a miserable learning experience.\nThis post is the practical follow-up the series has been missing. It covers the three things that make microservices debuggable: correlation IDs, distributed tracing, and centralized logs. And it shows you what the debugging experience actually looks like with and without them.\nWhat a stack trace used to look like In a monolith, when an order fails halfway through processing, you get something like this:\nERROR: failed to reserve stock for order 8842 at warehouse/reserve.go:47 at order/process.go:128 at order/create.go:63 at http/handler.go:22 Four lines. One process. You can read it top to bottom and see exactly where the error happened, what called it, and what the request was. The entire execution path lives in one place.\nWhat it looks like without preparation Now the same failure across three services, with no correlation ID:\nOrder service logs:\n2026-07-27T03:14:22Z INFO POST /orders 201 created order_id=8842 2026-07-27T03:14:22Z INFO calling warehouse service to reserve stock 2026-07-27T03:14:22Z ERROR POST http://warehouse:8080/reserve timeout after 5000ms Payment service logs:\n2026-07-27T03:14:23Z INFO POST /payments 201 created payment_id=10294 2026-07-27T03:14:23Z INFO stock reserved for order, proceeding with payment Warehouse service logs:\n2026-07-27T03:14:18Z INFO POST /reserve 200 stock reserved order_id=8842 Three services. Three log streams. Three different timestamps (the order service is three seconds ahead of the warehouse service; clock skew is real). You can\u0026rsquo;t tell which logs belong to the same request. You can\u0026rsquo;t tell that the payment service actually did reserve stock successfully, but the order service timed out waiting and retried, and the retry hit a race condition. You\u0026rsquo;re staring at three separate timelines that might or might not be related, and the customer is still waiting.\nThis is the single biggest operational difference between monoliths and microservices. In a monolith, the execution path is visible in one place. In microservices, reconstructing that path is a skill you have to build, and the tooling has to exist before you need it.\nCorrelation IDs: the minimum viable fix A correlation ID (sometimes called a request ID or trace ID) is a unique string generated when a request enters your system and passed along to every service that touches it. Every log line that service produces includes that ID. That\u0026rsquo;s it: one string that lets you grep across all services and reconstruct a single request\u0026rsquo;s journey.\nWhere to generate it At the entry point. When a request hits your API gateway or the first service in the chain, generate a UUID and attach it to the request:\nfunc middleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { id := r.Header.Get(\u0026#34;X-Request-ID\u0026#34;) if id == \u0026#34;\u0026#34; { id = uuid.New().String() } ctx := context.WithValue(r.Context(), \u0026#34;request_id\u0026#34;, id) next.ServeHTTP(w, r.WithContext(ctx)) }) } How to pass it Every outgoing HTTP or gRPC call from one service to another includes the correlation ID as a header:\nfunc callWarehouse(ctx context.Context, order Order) error { reqID := ctx.Value(\u0026#34;request_id\u0026#34;).(string) req, _ := http.NewRequestWithContext(ctx, \u0026#34;POST\u0026#34;, warehouseURL, body) req.Header.Set(\u0026#34;X-Request-ID\u0026#34;, reqID) // ... do the call } How to log it Every log line includes the correlation ID:\nlog.Printf(\u0026#34;[request_id=%s] stock reserved for order %d\u0026#34;, reqID, orderID) Now, when that same failure happens across three services, your grep looks like this:\n$ grep \u0026#34;request_id=abc-123\u0026#34; *.log order.log: INFO POST /orders 201 created request_id=abc-123 order.log: INFO calling warehouse service request_id=abc-123 order.log: ERROR POST http://warehouse:8080/reserve timeout request_id=abc-123 warehouse.log: INFO POST /reserve 200 stock reserved request_id=abc-123 payment.log: INFO POST /payments 201 created request_id=abc-123 You can now see the full timeline: the warehouse did reserve stock successfully, but the order service timed out before it got the response. The payment service then processed independently. You have a retry problem and a timeout configuration problem, not a missing request problem. Same incident, completely different debugging experience.\nDistributed tracing: what correlation IDs can\u0026rsquo;t do Correlation IDs solve the \u0026ldquo;which logs belong together\u0026rdquo; problem. They don\u0026rsquo;t solve the \u0026ldquo;how long did each step take\u0026rdquo; or \u0026ldquo;where did the time go\u0026rdquo; problem. For that, you need distributed tracing, and the practical entry point is OpenTelemetry.\nThe mental model A trace is a tree of spans. The root span is the incoming request. Each outbound call to another service creates a child span. Each span records its start time, end time, status, and metadata. When the trace is collected, you see a waterfall diagram showing exactly how time was spent:\n[POST /orders 4,200ms] ├─ [order.validate 12ms] ├─ [HTTP POST warehouse:/reserve 3,800ms] ← timeout at 5s but retried │ └─ [warehouse.reserve.stock 45ms] ├─ [HTTP POST payment:/charge 380ms] │ └─ [payment.process 350ms] └─ [order.confirm 8ms] That waterfall makes the timeout visible instantly. You can see the warehouse call took 3.8 seconds, the payment call took 380ms, and the total request took 4.2 seconds. Without tracing, you\u0026rsquo;d be guessing which step was slow.\nWhat it takes to add The basics aren\u0026rsquo;t as much infrastructure as you\u0026rsquo;d think. OpenTelemetry\u0026rsquo;s Go SDK gives you a tracer you add to your service\u0026rsquo;s entry point and each outbound call:\nimport \u0026#34;go.opentelemetry.io/otel\u0026#34; var tracer = otel.Tracer(\u0026#34;order-service\u0026#34;) func createOrder(ctx context.Context, req OrderRequest) (Order, error) { ctx, span := tracer.Start(ctx, \u0026#34;create-order\u0026#34;) defer span.End() // validate, call warehouse, call payment... // each outbound call gets its own child span } The SDK handles context propagation: when you make an outbound HTTP call, the trace context is automatically injected into the headers. The destination service picks it up and continues the trace. You don\u0026rsquo;t pass IDs manually; the instrumentation does it.\nWhat you do need is a collector: something that receives spans from all your services and stores them. Jaeger and Grafana Tempo are the common open-source options. The collector is the one piece of infrastructure you have to run, and it\u0026rsquo;s the thing that makes tracing data searchable.\nCorrelation IDs and distributed tracing aren\u0026rsquo;t alternatives; they serve different purposes. Correlation IDs are for grepping logs. Distributed tracing is for understanding latency. You want both, but you can start with correlation IDs and add tracing later. Don\u0026rsquo;t let the tracing infrastructure become the reason you don\u0026rsquo;t instrument at all.\nCentralized logs: why kubectl logs stops working With one or two services, reading logs from each one is tedious but feasible. With ten services, it\u0026rsquo;s impossible. You need your logs in one place (a system like Loki, the ELK stack, or Datadog) where you can search across all services at once.\nThe practical minimum: every service writes structured logs (JSON, not free-form text) to stdout, and a log collector ships them to a central store. Structured logs matter because they let you search by field (\u0026ldquo;show me all logs where request_id = abc-123\u0026rdquo;) instead of grepping free-form text and hoping the format is consistent.\nlog.Info(\u0026#34;stock reserved\u0026#34;, zap.String(\u0026#34;request_id\u0026#34;, reqID), zap.Int(\u0026#34;order_id\u0026#34;, orderID), zap.Int(\u0026#34;quantity\u0026#34;, qty), ) This isn\u0026rsquo;t glamorous work. It\u0026rsquo;s the kind of thing that feels unnecessary until you\u0026rsquo;re debugging a production incident and realize you can\u0026rsquo;t find the relevant logs because they\u0026rsquo;re scattered across fifteen containers and half of them are in a format you can\u0026rsquo;t search.\nThe debugging workflow: what it actually looks like With correlation IDs, tracing, and centralized logs in place, here\u0026rsquo;s what debugging a real incident looks like:\nCustomer reports the problem. \u0026ldquo;My order went through but I never got a shipping notification.\u0026rdquo;\nFind the request in your trace system. Search by order ID or customer email. You get the full trace: every service that handled the request, every span, every timing.\nSee where it broke. The trace shows Order → Warehouse (success) → Payment (success) → Notification (never called). The saga ended but didn\u0026rsquo;t trigger the final step.\nCheck the logs for context. Correlation ID in hand, search your centralized logs for every log line across all three services. You find the payment service logged a success but the event it published to the notification topic was never received, because the topic name was misspelled in the notification service\u0026rsquo;s config.\nFix it. Correct the topic name. Deploy. The trace confirms the notification span now completes.\nTotal debugging time: fifteen minutes. Without the tooling, that same incident could take hours: you\u0026rsquo;d be SSH-ing into containers, reading log files manually, and trying to reconstruct a timeline by hand.\nThe mistake to avoid: building the infrastructure after the incident The most common pattern I\u0026rsquo;ve seen on teams adopting microservices is this: they build the services, deploy them, ship features, and then when the first real incident happens, someone says \u0026ldquo;we should add tracing\u0026rdquo; and \u0026ldquo;we should centralize our logs.\u0026rdquo; At that point you\u0026rsquo;re debugging in the dark during your most stressful moment, and retrofitting observability under pressure is how you end up with inconsistent instrumentation that misses the important calls.\nThe minimum viable observability stack, before you deploy your first microservice:\nCorrelation ID middleware: twenty lines of code, generates an ID at the entry point, passes it on every outbound call, logs it on every log line. Structured logging: switch from free-form log.Printf to structured output with a library like zap or slog. Every log line includes the correlation ID and the key business identifiers. A log aggregator: Loki, ELK, or even a hosted solution. Something that lets you search across all services by correlation ID. Distributed tracing: OpenTelemetry with Jaeger or Tempo. Start with auto-instrumentation (the Go SDK instruments net/http automatically) and add manual spans for the business-critical paths. CPU and memory profiling: Go\u0026rsquo;s built-in pprof tool, exposed via a debug endpoint. When you\u0026rsquo;ve found which service is slow but not why, the profile tells you exactly which function is burning cycles. See Go performance profiling with pprof. None of this is optional infrastructure. It\u0026rsquo;s the equivalent of having a debugger in your language: you don\u0026rsquo;t ship a product without one, and you don\u0026rsquo;t ship microservices without observability.\nThis is the operational reality that design posts don\u0026rsquo;t cover: the part where your clean architecture has to survive contact with production. If you\u0026rsquo;re following the series from the beginning, the decision checklist told you whether you need microservices, the coupling and cohesion post told you where to draw the boundaries, and the communication and saga posts gave you the patterns. This post gives you the tooling to actually operate what you\u0026rsquo;ve built. Don\u0026rsquo;t skip it.\n","permalink":"https://hanhpham.vercel.app/posts/debugging-microservices-where-did-that-request-go/","summary":"In a monolith, a stack trace tells you what went wrong. In microservices, you\u0026rsquo;re grepping across three services\u0026rsquo; logs at 2 AM trying to figure out which one lost the request. Here\u0026rsquo;s how to make that possible instead of painful.","title":"Debugging Microservices: Where Did That Request Go?"},{"content":"MusicCorp\u0026rsquo;s warehouse service has an endpoint that returns stock levels:\n{ \u0026#34;sku\u0026#34;: \u0026#34;CD-NIRVANA-1991\u0026#34;, \u0026#34;quantity\u0026#34;: 42, \u0026#34;location\u0026#34;: \u0026#34;warehouse-3\u0026#34; } The product team decides that location should be a structured object instead of a flat string: warehouse-3 is in Portland, and downstream services need the city, not just the warehouse ID. The change is simple in the warehouse service. But six other services read that field. Three of them are owned by teams that deploy weekly. One of them is owned by a team that deploys quarterly. You can\u0026rsquo;t coordinate six simultaneous deploys. So the question becomes: how do you change location from a string to an object without breaking the three services that still expect a string?\nIn a monolith, this isn\u0026rsquo;t a question: you change the type, update every caller, and ship one commit. In microservices, you\u0026rsquo;re shipping a change that will run alongside the old version of your API for days or weeks while consumers upgrade at their own pace. Get this wrong and you get the most common microservices production incident: a service deploys a breaking change, and somewhere in the system, a consumer that hasn\u0026rsquo;t upgraded yet starts failing silently.\nWhy versioning is harder than you think In a monolith, the type system and the compiler catch breaking changes. If you rename a field, every place that references it fails to compile. You fix them all in one pass, ship one commit, and the old code stops existing the moment the new code deploys.\nIn microservices, there\u0026rsquo;s no compiler across services. When you deploy a change to Service A\u0026rsquo;s API, Services B through G don\u0026rsquo;t automatically update. They\u0026rsquo;re still running the old code that expects the old shape. Some of them might deploy today, some next week, some next quarter. During that window, the old and new versions of your API have to coexist; the old version has to keep working for consumers that haven\u0026rsquo;t upgraded yet.\nThis creates a compatibility matrix. With N services and M API versions, you might have N×M combinations to think about. In practice, you don\u0026rsquo;t, but only if you adopt a discipline that keeps the number of concurrent versions small and the transition path clear.\nThe expand-contract pattern The single most practical versioning strategy is expand-contract (also called parallel change or cross-phase). The rule is simple:\nExpand: add the new field alongside the old one. Both exist. Migrate: consumers switch to the new field at their own pace. Contract: remove the old field once all consumers have migrated. For the warehouse example:\nExpand: the API now returns both:\n{ \u0026#34;sku\u0026#34;: \u0026#34;CD-NIRVANA-1991\u0026#34;, \u0026#34;quantity\u0026#34;: 42, \u0026#34;location\u0026#34;: \u0026#34;warehouse-3\u0026#34;, \u0026#34;warehouse\u0026#34;: { \u0026#34;id\u0026#34;: \u0026#34;warehouse-3\u0026#34;, \u0026#34;city\u0026#34;: \u0026#34;Portland\u0026#34; } } Old consumers keep reading location. New consumers read warehouse.city. Nothing breaks. No coordination required.\nMigrate: each consuming team updates their code when they\u0026rsquo;re ready. Team A deploys on Tuesday. Team B deploys the following sprint. Team C, the quarterly deployers, gets a ticket in their backlog.\nContract: once you\u0026rsquo;ve confirmed (via logging or metrics) that no consumer is reading location anymore, remove it. The field is gone, the API is clean, and you\u0026rsquo;re back to one version.\nThe discipline is in the contract phase. Teams that expand but never contract end up with APIs full of deprecated fields that nobody dares remove. Add a metric: track how many requests still read the old field. When it drops to zero, remove it. When it doesn\u0026rsquo;t drop to zero, find out who\u0026rsquo;s still reading it and talk to them.\nURL versioning vs. header versioning When you need genuinely incompatible changes (not just adding a field but changing the semantics of the entire response), you need to version the API itself. There are two practical approaches:\nURL versioning (/v1/stock, /v2/stock) is explicit and visible. You can see which version you\u0026rsquo;re calling. It\u0026rsquo;s easy to route, easy to document, and easy to debug. The downside: it leaks implementation details into your URLs, and it gives consumers an excuse to stay on v1 forever because v1 still works.\nHeader versioning (Accept: application/vnd.musiccorp.v2+json) keeps your URLs clean and makes versioning a negotiation between client and server. The downside: it\u0026rsquo;s invisible in logs, harder to test with curl, and easier to forget to set.\nThe pragmatic choice for most teams: URL versioning for major breaking changes, expand-contract for everything else. If you\u0026rsquo;re versioning your API more than twice, something is wrong with your evolution discipline: you\u0026rsquo;re making breaking changes too often.\nSchema evolution: protobuf and Avro If you\u0026rsquo;re using protocol buffers or Avro for inter-service communication instead of JSON, you get versioning tools built into the serialization format. Protobuf\u0026rsquo;s field numbering system is designed for exactly this problem:\nmessage StockLevel { string sku = 1; int32 quantity = 2; string location = 3; // deprecated, but still readable Warehouse warehouse = 4; // the replacement } message Warehouse { string id = 1; string city = 2; } Protobuf\u0026rsquo;s rule: old code ignores unknown fields. So when you add field 4, old consumers that don\u0026rsquo;t know about it simply skip it: no error, no crash. When you remove field 3, new consumers that expected it check for its presence. This gives you safe forward and backward compatibility by default, as long as you follow two rules:\nNever reuse a field number. If you remove field 3, leave it reserved. Don\u0026rsquo;t assign a new meaning to number 3; old code might still be sending data with that number. Use optional for fields that might not be present. New consumers checking a removed field need to handle its absence. Avro gives you something even better for schema evolution: a writer\u0026rsquo;s schema and a reader\u0026rsquo;s schema that the system reconciles automatically. If the writer sends field A and field B, and the reader expects field B and field C`, the system fills in a default for C and ignores A. You don\u0026rsquo;t have to think about forward compatibility per se; the format handles it.\nIf you\u0026rsquo;re building new services and choosing a wire format, protobuf or Avro over JSON is worth the upfront investment purely for the versioning story. JSON is human-readable and easy to prototype with, but it gives you zero help with schema evolution: every field change is a potential breaking change that only your tests (if you have them) will catch.\nConsumer-driven contracts Expand-contract handles the \u0026ldquo;add then remove\u0026rdquo; lifecycle. But how do you know when it\u0026rsquo;s safe to remove the old field? How do you know that no consumer still depends on it?\nConsumer-driven contracts flip the testing direction. Instead of the provider testing that its API works, each consumer defines what it needs: \u0026ldquo;I expect GET /stock/{sku} to return an object with quantity as an integer.\u0026rdquo; The provider runs all consumers\u0026rsquo; contracts against its API before deploying. If any contract fails, the deploy is blocked.\nThe practical entry point is Pact, an open-source contract testing framework. The consumer writes a contract:\n# consumer side def test_stock_level_has_quantity(provider): provider.given(\u0026#34;sku CD-NIRVANA-1991 exists\u0026#34;) provider.upon_receiving(\u0026#34;a request for stock level\u0026#34;) provider.with_request(\u0026#34;GET\u0026#34;, \u0026#34;/stock/CD-NIRVANA-1991\u0026#34;) provider.will_respond_with(200, body={ \u0026#34;quantity\u0026#34;: Like(42) }) The provider verifies this contract against its real API. If the provider tries to remove quantity and a consumer\u0026rsquo;s contract still expects it, the test fails and the deploy is stopped.\nThis is the tooling that makes the contract phase of expand-contract reliable. Without it, you\u0026rsquo;re guessing whether anyone still reads the old field. With it, you know.\nThe database schema problem Versioning an API is tractable. Versioning a database schema that multiple services share is where things get genuinely hard, and it\u0026rsquo;s the one scenario where microservices can trap you if you\u0026rsquo;re not careful.\nIf Service A and Service B both read from the same stock_levels table, and Service A decides to rename the quantity column to available_count, Service B breaks immediately. Not on the next deploy: right now, the moment the migration runs.\nThe fix is architectural: each service owns its database. If Service B needs stock data, it gets it through Service A\u0026rsquo;s API, not by querying the table directly. This is the rule from the coupling post: no shared databases.\nBut if you\u0026rsquo;re mid-migration and the shared database is real, the practical stopgap is:\nAdd the new column first (expand), write to both during a transition period. Update consumers to read from the new column. Drop the old column once reads have migrated. It\u0026rsquo;s expand-contract at the database level. The danger is that database migrations are harder to roll back than API changes; dropping a column isn\u0026rsquo;t undone by redeploying the old version of your service. Test migrations against a copy of production data before running them.\nWhat it looks like when you get it wrong Here\u0026rsquo;s the incident. MusicCorp\u0026rsquo;s order service sends stock reservation requests to the warehouse service. The request includes a priority field as a string: \u0026quot;normal\u0026quot; or \u0026quot;express\u0026quot;.\nThe warehouse team decides to change priority from a string to an enum integer: 0 for normal, 1 for express. They deploy the change with no expand phase: the old string format is gone, only the new integer format is accepted.\nThe order service, which still sends \u0026quot;priority\u0026quot;: \u0026quot;normal\u0026quot;, starts failing. The warehouse returns a 400 Bad Request. The order service retries, gets 400 again, retries three times, and then the order fails. The customer sees an error page.\nBut here\u0026rsquo;s the worse part: it only happens for express orders. Normal orders happen to work because the warehouse service\u0026rsquo;s validation has a bug that treats the missing priority field as normal priority. So the failure is intermittent, tied to a specific code path, and the error message is a generic 400 with no hint that it\u0026rsquo;s a type mismatch.\nTotal time to identify: four hours, because the first investigation focused on the order service (which was retrying) rather than the warehouse service (which was rejecting). A correlation ID and a trace would have shown the 400 immediately. The expand-contract pattern would have prevented it entirely.\nVersioning is the boring discipline that keeps microservices from rotting into a system where every deploy is a coordination exercise. The expand- contract pattern is your default strategy, protobuf or Avro gives you safe schema evolution for free, and consumer-driven contracts tell you when it\u0026rsquo;s safe to remove the old version. None of this is exciting. All of it is necessary.\nIf you\u0026rsquo;re following the series, the communication tech post covered how to pick the wire format; this post covers how to evolve it safely. And if the incident scenario above made you wince, the debugging post covers the tooling that would have cut the investigation time from four hours to fifteen minutes.\n","permalink":"https://hanhpham.vercel.app/posts/microservices-versioning-without-breaking-everyone/","summary":"In a monolith, renaming a field means updating all callers in one commit. In microservices, consumers upgrade on their own schedule, so your API has to evolve without breaking the ones that haven\u0026rsquo;t upgraded yet. Here\u0026rsquo;s how.","title":"Versioning Microservices Without Breaking Everyone"},{"content":"The decision checklist mentioned this in passing: running a monolith locally means go run . and you\u0026rsquo;re done. Running microservices locally means Docker Compose with fifteen containers, or a shared staging environment that\u0026rsquo;s always half-broken. This post is the practical version of that warning: how to actually structure local development so it doesn\u0026rsquo;t become the tax that makes everyone wish you\u0026rsquo;d stayed with the monolith.\nThe core problem is simple: a microservice architecture assumes network calls between services. On your laptop, you don\u0026rsquo;t have a network with twelve running services. You have one machine, limited RAM, and a Docker daemon that slows to a crawl once you pass six or seven containers. The question is how to get the service you\u0026rsquo;re working on running locally without needing every other service to be running too.\nThe Docker Compose default, and why it breaks down The first instinct is a docker-compose.yml that runs everything:\nservices: order: build: ./order-service warehouse: build: ./warehouse-service payment: build: ./payment-service notification: build: ./notification-service inventory: build: ./inventory-service # ... seven more services This works on day one. By day thirty, it doesn\u0026rsquo;t. Here\u0026rsquo;s why:\nRAM. Each container runs a full runtime. A Go service might use 30MB. A Java service uses 256MB minimum. A PostgreSQL container uses 100MB. A Redis container uses 50MB. Twelve services plus infrastructure and you\u0026rsquo;re at 2–3GB before your IDE loads. On an 8GB laptop, you\u0026rsquo;re swapping.\nStartup time. docker compose up starts services in dependency order. Service A waits for Service B, which waits for Service C. By the time everything is healthy, you\u0026rsquo;ve made coffee, checked Slack, and forgotten what you were working on.\nFragility. Service D was deployed yesterday with a new environment variable. Your local docker-compose.yml doesn\u0026rsquo;t set it. Service D crashes. Service A, which depends on Service D, also crashes. You spend forty minutes figuring out which container broke and why, and the actual change you were making (a one-line fix in Service B) has nothing to do with any of it.\nThe shared docker-compose.yml is the most common way teams try to solve local development for microservices, and it\u0026rsquo;s the one that scales worst. It works for two or three services. Past that, it becomes a maintenance burden that fights you more than it helps.\nPattern 1: run only what you\u0026rsquo;re working on The simplest fix: don\u0026rsquo;t run everything. Run the service you\u0026rsquo;re changing, and mock the ones it talks to.\nIf you\u0026rsquo;re working on the order service, you need:\nThe order service itself (running from your IDE, not Docker) A database for the order service (one PostgreSQL container) Mocks for the warehouse, payment, and notification services That\u0026rsquo;s four things instead of twelve. It fits in 500MB of RAM. It starts in seconds. And the mocks can be simple HTTP servers that return canned responses:\n// mock_server.go: run this, forget about it func main() { http.HandleFunc(\u0026#34;/reserve\u0026#34;, func(w http.ResponseWriter, r *http.Request) { w.WriteHeader(http.StatusOK) json.NewEncoder(w).Encode(map[string]interface{}{ \u0026#34;status\u0026#34;: \u0026#34;reserved\u0026#34;, \u0026#34;order_id\u0026#34;: \u0026#34;mock-123\u0026#34;, }) }) http.ListenAndServe(\u0026#34;:8081\u0026#34;, nil) } This isn\u0026rsquo;t production-realistic, but it doesn\u0026rsquo;t need to be. You\u0026rsquo;re developing the order service\u0026rsquo;s logic, not the warehouse service\u0026rsquo;s behavior. The mock tells you \u0026ldquo;the warehouse accepted the request\u0026rdquo; so you can verify that the order service handles the response correctly.\nPattern 2: service virtualization Canned mocks work, but they don\u0026rsquo;t handle edge cases: what happens when the warehouse returns a 503? What happens when the response is missing a field? For that, you need a mock that can be configured per-test:\nWireMock (for HTTP) lets you define stub mappings:\n{ \u0026#34;request\u0026#34;: { \u0026#34;method\u0026#34;: \u0026#34;POST\u0026#34;, \u0026#34;url\u0026#34;: \u0026#34;/reserve\u0026#34; }, \u0026#34;response\u0026#34;: { \u0026#34;status\u0026#34;: 503, \u0026#34;body\u0026#34;: \u0026#34;{\\\u0026#34;error\\\u0026#34;: \\\u0026#34;service unavailable\\\u0026#34;}\u0026#34; }, \u0026#34;priority\u0026#34;: 10 } You can set different responses for different request bodies, add delays, or simulate flaky behavior. The mock behaves like the real service without running it.\nMountebank does the same thing across protocols: HTTP, TCP, and amqp. If your services communicate over a message broker, Mountebank can mock the broker\u0026rsquo;s behavior.\nThe investment is real (you have to write and maintain the stubs), but it\u0026rsquo;s a one-time cost per service boundary, and it pays for itself every time someone can reproduce a bug locally instead of deploying to a shared staging environment and waiting.\nPattern 3: Docker Compose profiles If you do need some real services running (maybe you\u0026rsquo;re testing a workflow that spans three services and mocking isn\u0026rsquo;t realistic), Docker Compose profiles let you choose which services to start:\nservices: order: build: ./order-service profiles: [\u0026#34;full\u0026#34;] warehouse: build: ./warehouse-service profiles: [\u0026#34;full\u0026#34;, \u0026#34;fulfillment\u0026#34;] payment: build: ./payment-service profiles: [\u0026#34;full\u0026#34;, \u0026#34;fulfillment\u0026#34;] postgres: image: postgres:16 # no profile: always runs # just the database: for working on order service logic docker compose up postgres # order + warehouse + payment: for testing fulfillment docker compose --profile fulfillment up # everything: for integration tests docker compose --profile full up This doesn\u0026rsquo;t solve the RAM problem, but it solves the \u0026ldquo;I only need three services, not twelve\u0026rdquo; problem. Profiles make the compose file a menu instead of an all-or-nothing commitment.\nPattern 4: remote development environments The nuclear option that actually works: run the services somewhere else.\nTilt and Skaffold watch your local code, rebuild the container image on change, and deploy to a local Kubernetes cluster (like minikube or k3d) or a remote cluster. You edit code locally, Tilt syncs it to the cluster, and the cluster runs the full architecture with proper networking.\nGitpod and GitHub Codespaces give you a cloud VM with the full stack pre-configured. New developers clone the repo, open the IDE, and everything is already running. No Docker setup, no port conflicts, no \u0026ldquo;works on my machine.\u0026rdquo;\nTelepresence lets you run one service locally while proxying to a remote cluster for everything else. You get hot-reload on the service you\u0026rsquo;re changing and real dependencies for everything else.\nThese solutions cost money (cloud VMs, cluster resources) and add complexity (you need to maintain the remote environment). But for teams of five or more, the time saved per developer per week often justifies it quickly. The question is whether the operational cost is less than the cumulative cost of every developer fighting Docker Compose every morning.\nThe shared staging trap When local development is painful, teams default to a shared staging environment: \u0026ldquo;just test your changes on staging, it has all the services running.\u0026rdquo; This is the trap.\nShared staging breaks constantly because:\nEveryone\u0026rsquo;s changes are deployed to the same environment simultaneously Database state is shared: one person\u0026rsquo;s test data breaks another person\u0026rsquo;s test The environment drifts from production because nobody maintains it with the same discipline Debugging failures requires asking \u0026ldquo;was this my change or someone else\u0026rsquo;s?\u0026rdquo; Shared staging is useful for final validation before production: running the full integration suite against a realistic environment. It\u0026rsquo;s not useful for development: writing and testing your code before you ship it. If your workflow is \u0026ldquo;write code, push to staging, see if it works, fix, repeat,\u0026rdquo; you\u0026rsquo;ve turned a local debugging problem into a shared-environment debugging problem, which is worse.\nThe rule of thumb: if you can\u0026rsquo;t test your change without deploying to a shared environment, your local development setup isn\u0026rsquo;t complete. You need to be able to run the service you\u0026rsquo;re changing with enough of its dependencies to verify its behavior (real or mocked) before anyone else sees it.\nWhat actually works in practice After seeing several teams go through this, the pattern that scales is a layered approach:\nUnit tests run locally, no infrastructure. The service\u0026rsquo;s business logic is tested with mocks for external dependencies. This is your fastest feedback loop: seconds, not minutes.\nIntegration tests run with Docker Compose profiles. The service plus its direct dependencies (database, one or two adjacent services) run in containers. This takes a minute to start and tests the real wiring.\nContract tests verify boundaries. Consumer-driven contracts (from the versioning post) ensure your service\u0026rsquo;s API matches what consumers expect, without running the consumers.\nEnd-to-end tests run on a ephemeral environment. A fresh copy of the full stack, spun up for the test suite and torn down after. Not shared, not persistent, not a place where state accumulates.\nEach layer catches a different class of bug. Unit tests catch logic errors. Integration tests catch wiring errors. Contract tests catch boundary misunderstandings. End-to-end tests catch workflow failures. No single layer is sufficient, and no layer should be skipped, but the earlier layers should catch most bugs before you need the expensive ones.\nLocal development is the tax you pay every day for the independent deployability that microservices give you. The teams that handle it well are the ones that invest in the tooling early (mocks, profiles, contract tests) instead of discovering the pain in their first week and spending the next three months fighting Docker Compose.\n","permalink":"https://hanhpham.vercel.app/posts/microservices-local-dev-without-docker-compose-hell/","summary":"You split the monolith into twelve services. Now a new developer\u0026rsquo;s first week is spent getting Docker Compose to run all twelve on their laptop, and half of them crash because of a port conflict. Here\u0026rsquo;s how to make local development work without the pain.","title":"Running Microservices Locally Without Docker Compose Hell"},{"content":"The decision checklist listed \u0026ldquo;independent deployability\u0026rdquo; as the core value proposition of microservices. Team A ships on Tuesday without waiting for Team B. That\u0026rsquo;s the promise. The reality is that independent deployability requires a deployment strategy, a CI/CD pipeline per service, a rollback plan, and a way to know whether your deploy just broke a consumer that hasn\u0026rsquo;t upgraded yet. Without those, \u0026ldquo;independent deployability\u0026rdquo; is just \u0026ldquo;everyone deploys whenever they want and hopes for the best.\u0026rdquo;\nThis post covers the deployment patterns that make independent deploys actually work (rolling deploys, blue-green, canary, and feature flags), plus the CI/CD structure that makes them repeatable. It\u0026rsquo;s the operational follow-up to the versioning post, which covered how to evolve APIs without breaking consumers. This post covers how to get the new code running without downtime.\nThe deploy shift: from one thing to N things In a monolith, a deploy is simple:\nBuild the artifact. Stop the old instance. Start the new instance. Verify it\u0026rsquo;s healthy. One thing to build, one thing to stop, one thing to start. The entire deploy is atomic: either the old version is running or the new version is running, never both.\nIn microservices, a deploy is:\nBuild the artifact for Service A. Deploy Service A while Services B through L keep running. Verify Service A is healthy. Verify that Services B through L still work with the new version of A. Step 4 is the one that doesn\u0026rsquo;t exist in a monolith. You\u0026rsquo;ve changed a service that other services depend on. Those services are still running the old code. The new version of A has to work with the old versions of its consumers, and the old version of A (if it\u0026rsquo;s still running during a rolling deploy) has to work with consumers that might have already upgraded. This is the versioning problem in deployment form.\nRolling deploys: the default A rolling deploy replaces instances one at a time. At any point during the deploy, some instances run the old version and some run the new. This is the simplest strategy and the one most teams start with:\nBefore: [v1] [v1] [v1] [v1] Deploy: [v2] [v1] [v1] [v1] ← first instance updated [v2] [v2] [v1] [v1] ← second instance updated [v2] [v2] [v2] [v1] ← third instance updated After: [v2] [v2] [v2] [v2] ← all instances updated Rolling deploys require that both versions can run simultaneously, which means the API must be backward-compatible during the deploy window. This is exactly the expand phase of expand-contract: the new version adds something, the old version still works, and once all instances are updated, you can contract.\nThe risk: if the new version has a bug, it\u0026rsquo;s rolling out to production gradually. You\u0026rsquo;ll see errors accumulate as more instances switch over. The fix is health checks: if the new version fails its health check, the deploy stops and rolls back automatically.\nBlue-green deploys: instant rollback A blue-green deploy runs two identical environments: \u0026ldquo;blue\u0026rdquo; (current) and \u0026ldquo;green\u0026rdquo; (new). You deploy to green, verify it works, then switch the load balancer from blue to green. If something breaks, you switch back:\nBefore: traffic → [blue: v1] Deploy: traffic → [blue: v1] green: [v2] ← deploying Verify: traffic → [blue: v1] green: [v2] ← health checks pass Switch: traffic → [green: v2] blue: [v1] ← live Rollback: traffic → [blue: v1] green: [v2] ← instant switch back The advantage: rollback is instant. You don\u0026rsquo;t have to rebuild and redeploy the old version; it\u0026rsquo;s still running in blue. The cost: you need twice the infrastructure during the deploy. For a service running four instances, you need eight during the switch. For a small team with limited infrastructure, this might be too expensive. For a service handling critical traffic, it\u0026rsquo;s worth it.\nCanary deploys: test with real traffic A canary deploy sends a small percentage of traffic to the new version while the rest stays on the old:\nBefore: traffic → [v1] [v1] [v1] [v1] Canary: 5% traffic → [v2] 95% traffic → [v1] [v1] [v1] [v1] Full: traffic → [v2] [v2] [v2] [v2] This is the safest deploy strategy because you\u0026rsquo;re testing the new version with real production traffic before committing to it. If the canary shows increased error rates, latency spikes, or business metric anomalies, you kill it and stay on v1.\nThe tooling for this is more complex: you need a load balancer or service mesh that can split traffic by percentage, and you need metrics collection that can compare canary vs. baseline in real time. Istio and Linkerd do this natively. If you\u0026rsquo;re not running a service mesh, most cloud load balancers (ALB, Cloudflare) support weighted target groups.\nThe practical entry point: start with rolling deploys and health checks. Move to blue-green when the cost of downtime justifies the infrastructure cost. Move to canary when you have the metrics and traffic-splitting infrastructure to make it meaningful. Don\u0026rsquo;t start with canary; it\u0026rsquo;s the most sophisticated strategy and the hardest to get right.\nFeature flags: deploy without releasing Deploying code and releasing features are two different things. A feature flag lets you deploy code that\u0026rsquo;s off by default, then turn it on incrementally:\nif featureflags.Enabled(ctx, \u0026#34;new-checkout-flow\u0026#34;) { return newCheckout(ctx, order) } return legacyCheckout(ctx, order) Deploy the code with the flag off. The new code path exists in production but isn\u0026rsquo;t executed. Turn it on for 1% of users, then 10%, then 100%. If something breaks, turn the flag off; no redeploy needed.\nFeature flags solve a problem that deploy strategies alone don\u0026rsquo;t: the ability to separate \u0026ldquo;code is running in production\u0026rdquo; from \u0026ldquo;users can see this.\u0026rdquo; You can deploy on Friday (because the deploy is safe, the flag is off) and release on Monday (because you turn the flag on during business hours).\nThe danger: flag debt. Every feature flag is a branch in your code that someone has to maintain, test, and eventually remove. Teams that add flags aggressively without a removal process end up with a codebase full of dead code paths that nobody understands. The discipline: every flag gets a ticket to remove it. If the flag has been on for two weeks with no issues, remove the flag and the old code path.\nCI/CD structure: pipeline per service Every service needs its own CI/CD pipeline. The pipeline should:\nRun unit tests: fast, no infrastructure, catches logic bugs. Run integration tests: the service plus its direct dependencies (see the local dev post for how to set these up). Build the container image: tagged with the commit SHA, not latest. Run contract tests: verify the service\u0026rsquo;s API matches what consumers expect (from the versioning post). Deploy to staging: a fresh environment, not a shared one that accumulates state. Run smoke tests: a handful of end-to-end tests against staging. Deploy to production: with the strategy of your choice. The pipeline is per-service because each service deploys independently. Service A\u0026rsquo;s pipeline runs when Service A changes. Service B\u0026rsquo;s pipeline doesn\u0026rsquo;t run. That\u0026rsquo;s the whole point.\nIf you\u0026rsquo;re using a monorepo (all services in one repository), the pipeline needs to detect which services changed and only build/deploy those. Tools like Bazel, Turborepo, and nx handle this. Without change detection, a monorepo CI pipeline rebuilds everything on every commit, which defeats the purpose of independent deploys.\nWhat it looks like when a deploy breaks consumers MusicCorp\u0026rsquo;s warehouse service deploys a new version. The API changes quantity from an integer to a string (someone thought it should be \u0026ldquo;42\u0026rdquo; instead of 42; don\u0026rsquo;t ask why). The deploy succeeds: the warehouse service starts, passes health checks, and begins accepting requests.\nThe order service, which still sends {\u0026quot;quantity\u0026quot;: 42}, starts getting 400 errors from the warehouse. The order service retries, gets more 400 errors, and eventually fails the order. Customers see errors.\nThe warehouse team\u0026rsquo;s dashboard shows green: their service is healthy, the deploy succeeded, error rates are zero (because the warehouse service doesn\u0026rsquo;t count client errors as its own failures). The order team\u0026rsquo;s dashboard shows red: their error rate just spiked. Neither team can see the full picture.\nA deploy webhook (a notification sent to a shared channel when any service deploys) would have told the order team \u0026ldquo;warehouse just deployed\u0026rdquo; the moment it happened. An error budget (a shared metric that tracks total system health, not per-service health) would have shown the impact immediately. And the correlation ID from the debugging post would have connected the order service\u0026rsquo;s 400s to the warehouse service\u0026rsquo;s deploy in the trace.\nThe deploy discipline Independent deployability is a capability, not a default. It requires:\nBackward-compatible API changes: the expand-contract pattern. Health checks: every service exposes a /health endpoint that verifies database connectivity, downstream dependencies, and critical paths. The deploy orchestrator (Kubernetes, ECS, whatever) uses this to decide whether a new instance is ready to receive traffic. Automated rollback: if the health check fails during a deploy, the orchestrator reverts to the previous version automatically. No human intervention required. Deploy notifications: every deploy posts to a shared channel with the service name, version, and who triggered it. When something breaks, the first question is always \u0026ldquo;did anyone deploy recently?\u0026rdquo; Canary analysis: for critical services, compare error rates and latency between the canary and baseline before promoting the deploy. None of this is optional infrastructure. It\u0026rsquo;s the price of independent deployability: the thing that\u0026rsquo;s supposed to make microservices worth the operational cost. If you\u0026rsquo;re not willing to invest in deploy tooling, you\u0026rsquo;re not ready for microservices, because the alternative is coordinated deploys, which is what you were trying to escape.\nThe series now covers the full lifecycle: deciding whether you need microservices, drawing the boundaries, choosing how services communicate, keeping APIs backward-compatible, running the stack locally, debugging when things go wrong, and deploying without fear. Each post stands alone, but together they\u0026rsquo;re the playbook I wish someone had handed me before going through this the first time.\n","permalink":"https://hanhpham.vercel.app/posts/microservices-deploying-without-the-fear/","summary":"In a monolith, deploy means pushing one thing. In microservices, every service deploys independently, which is the whole point, until you realize that \u0026lsquo;independently\u0026rsquo; also means \u0026lsquo;without knowing what the other services are doing right now.\u0026rsquo;","title":"Deploying Microservices Without the Fear"},{"content":"The sagas post covered why services shouldn\u0026rsquo;t share database transactions. This post covers the architectural consequence of that rule: if each service owns its own database, how do you answer questions that span multiple services?\nMusicCorp\u0026rsquo;s order service knows about orders. The payment service knows about payments. The warehouse service knows about stock. The customer service knows about addresses and preferences. When a customer asks \u0026ldquo;show me my recent orders with their payment status and shipping tracking,\u0026rdquo; that query touches four databases. In a monolith, it\u0026rsquo;s one SQL join. In microservices, there\u0026rsquo;s no join across service boundaries, and that\u0026rsquo;s the point. The question is what you do instead.\nWhy database-per-service is the rule Before the patterns, the reason: if two services share a database, a schema change in one service breaks the other. Service A renames a column. Service B, which queries that column directly, starts failing. You\u0026rsquo;ve reintroduced the coupling you were trying to escape, not through the API, but through the data layer.\nWorse, shared databases prevent independent evolution. Service A can\u0026rsquo;t change its data model without coordinating with every other service that reads from the same tables. You end up with a de facto API that\u0026rsquo;s the database schema, and changing it requires the same coordination you\u0026rsquo;d need for an API change, except there\u0026rsquo;s no versioning, no expand-contract, and no documentation. It\u0026rsquo;s the worst of both worlds.\nThe rule exists for the same reason the coupling post exists: to keep service boundaries meaningful. A service that owns its database can evolve its data model freely, run migrations without coordinating with other teams, and scale its storage independently. That\u0026rsquo;s the value proposition. The cost is the query problem.\nPattern 1: API composition The simplest approach: the service that needs data from other services calls their APIs and composes the result.\nfunc GetOrderSummary(ctx context.Context, orderID string) (OrderSummary, error) { order, err := orderClient.Get(ctx, orderID) if err != nil { return OrderSummary{}, err } payment, err := paymentClient.GetByOrder(ctx, orderID) if err != nil { return OrderSummary{}, err } shipment, err := warehouseClient.GetShipment(ctx, orderID) if err != nil { return OrderSummary{}, err } return OrderSummary{ Order: order, Payment: payment, Shipment: shipment, }, nil } This is a composition query: one service acts as the aggregator, calls multiple downstream services, and assembles the result. It\u0026rsquo;s straightforward and easy to understand.\nThe problems:\nLatency adds up. Three sequential calls at 50ms each = 150ms minimum. You can parallelize with goroutines, but you still wait for the slowest service. Partial failures are tricky. What if the warehouse service is down? Do you return the order and payment data without shipment info? Return an error? The answer depends on the use case, and you have to decide for every composition query. No cross-service joins. You can\u0026rsquo;t sort by \u0026ldquo;orders with payments over $100, shipped in the last week\u0026rdquo; without fetching everything and filtering in memory. That works for small result sets. It doesn\u0026rsquo;t work for analytics. API composition is the right choice for simple read paths where the number of downstream calls is small (two or three) and the result set is bounded (a single order, a user profile, a product detail page).\nPattern 2: CQRS (separate reads from writes) Command Query Responsibility Segregation splits the data model into two: a write model (the service\u0026rsquo;s authoritative database, optimized for transactions) and a read model (a separate store optimized for queries).\nThe write side stays as-is: the order service has its orders table, the payment service has its payments table, each with its own schema optimized for writes. The read side is different: a read-optimized store that contains pre-joined, denormalized data assembled from multiple services' write models.\nWrite side: Read side: ┌─────────────┐ ┌──────────────────────┐ │ Order DB │──events──→ │ Order Summary Store │ │ Payment DB │──events──→ │ (denormalized view) │ │ Warehouse DB│──events──→ │ │ └─────────────┘ └──────────────────────┘ ↑ Read queries go here When an order is placed, the order service writes to its database and publishes an event. A projection service consumes that event, joins it with payment and shipment data, and writes a denormalized \u0026ldquo;order summary\u0026rdquo; row to the read store. When the customer asks for their order history, the query hits the read store: one fast SELECT, no cross-service calls.\nThe read store can be anything: a PostgreSQL table with the right indexes, an Elasticsearch index for full-text search, a Redis cache for hot data. The point is that it\u0026rsquo;s shaped for the query, not for the write.\nThe trade-off is eventual consistency. The read model is updated asynchronously. There\u0026rsquo;s a window (typically milliseconds) between \u0026ldquo;the write happened\u0026rdquo; and \u0026ldquo;the read model reflects it.\u0026rdquo; If the customer places an order and immediately refreshes the page, they might not see it yet. For most use cases, this is fine. For financial reporting or audit trails, it might not be.\nCQRS is the right choice when:\nYou have complex read patterns that span multiple services. Read performance matters more than read freshness. You\u0026rsquo;re willing to maintain a projection service and a read store. Pattern 3: Event sourcing (store what happened, not what is) Event sourcing takes the CQRS write model further: instead of storing the current state of an entity, you store every event that led to it.\nOrder 8842 events: 1. OrderPlaced {sku: \u0026#34;CD-NIRVANA-1991\u0026#34;, qty: 1, customer: \u0026#34;c-42\u0026#34;} 2. PaymentTaken {amount: 24.99, method: \u0026#34;visa\u0026#34;} 3. StockReserved {warehouse: \u0026#34;warehouse-3\u0026#34;} 4. OrderShipped {tracking: \u0026#34;1Z999AA10123456784\u0026#34;} The current state is derived by replaying events: start with an empty order, apply each event in order, and you end up with the current state. This gives you a complete audit trail for free: you can reconstruct the state at any point in time, not just the current state.\nEvent sourcing solves the cross-service query problem in a different way: events are the shared language. The order service emits OrderPlaced. The payment service consumes it and emits PaymentTaken. The warehouse service consumes OrderPlaced and emits StockReserved. A read model (following CQRS) consumes all three and builds the denormalized view.\nThe cost:\nComplexity. Event sourcing is conceptually simple but operationally complex. Event schemas evolve. You need snapshotting to avoid replaying thousands of events. You need to handle event versioning. Debugging is different. You can\u0026rsquo;t look at a row in a table and see the current state; you have to replay events to derive it. Tooling helps, but it\u0026rsquo;s a different mental model. It\u0026rsquo;s a bigger commitment than CQRS. CQRS separates reads from writes; event sourcing changes how you store writes entirely. You can do CQRS without event sourcing. You can\u0026rsquo;t easily do event sourcing without CQRS. Event sourcing is the right choice when:\nYou need a complete audit trail (finance, healthcare, legal). The domain is naturally event-driven (order lifecycle, IoT telemetry, collaborative editing). You want temporal queries (\u0026ldquo;what was the state of this order on July 1st?\u0026rdquo;). A common mistake: adopting event sourcing for the whole system when only one aggregate needs it. Start with CQRS for the query problem. Add event sourcing to specific aggregates where the audit trail or temporal query justifies the complexity. Don\u0026rsquo;t event-source your user preferences service.\nPattern 4: the reporting database For analytics and business intelligence, none of the above patterns are quite right. You don\u0026rsquo;t want to compose API calls in real time for a dashboard that queries six months of data. You want a reporting database: a copy of all relevant data, assembled into a single store, optimized for analytical queries.\nThe mechanism is the same as CQRS projections: services emit events, a pipeline consumes them, and a reporting database stores the denormalized result. But the reporting database isn\u0026rsquo;t serving real-time reads; it\u0026rsquo;s serving batch queries that run across the entire dataset.\nOrder events ──→ Payment events ──→ ETL pipeline ──→ Reporting DB (analytics) Warehouse events ──→ This is the one place where a shared data store makes sense, but it\u0026rsquo;s a copy, not the source of truth. The reporting database doesn\u0026rsquo;t serve writes. It doesn\u0026rsquo;t influence business logic. It\u0026rsquo;s a read-only view assembled from events, and if it falls behind, the reports are slightly stale but the system keeps running.\nThe tooling varies: Apache Kafka Connect can stream events to a data warehouse. Debezium can capture database changes and replicate them to a central store. For simpler setups, a scheduled job that calls each service\u0026rsquo;s API and writes the results to a PostgreSQL database works fine.\nWhen to break the rule The \u0026ldquo;database per service\u0026rdquo; rule exists to prevent coupling. But there are cases where a shared database is the pragmatic choice:\nSmall systems with one team. If three services are owned by the same team and deploy together, the coupling cost of a shared database is low and the query benefit is high. Don\u0026rsquo;t pay the CQRS tax if you don\u0026rsquo;t need to. Read-heavy, write-light data. If a service mostly reads data that another service owns (like a product catalog), direct database access is simpler than an API composition layer, as long as the schema is stable and the ownership is clear. Migration in progress. If you\u0026rsquo;re splitting a monolith and haven\u0026rsquo;t finished extracting services, a shared database is a temporary reality. The goal is to make it temporary: add a ticket to finish the extraction, and treat every shared table as technical debt. The test: if you can change the schema without coordinating with another team, the database is yours. If you can\u0026rsquo;t, you\u0026rsquo;re sharing, and you\u0026rsquo;re paying the coupling tax, whether you call it a shared database or not.\nThe sagas post covered how to keep writes consistent across services. This post covers how to keep reads performant. Together, they answer the full data question: writes use sagas for coordination, reads use composition, CQRS, or event sourcing depending on complexity and freshness requirements. The reporting database covers analytics. And sometimes, the pragmatic answer is a shared database with clear ownership; just make sure the coupling is a conscious choice, not an accident.\n","permalink":"https://hanhpham.vercel.app/posts/microservices-database-per-service/","summary":"Each service owns its database. That\u0026rsquo;s the rule. But when a report needs data from four services, or a customer wants to see their complete order history, you\u0026rsquo;ve got a query problem that a single SELECT can\u0026rsquo;t solve. Here\u0026rsquo;s how to handle it.","title":"Database Per Service: The Pattern That's Harder Than It Sounds"},{"content":"Every few months, someone on the team reads about service meshes and proposes adding Istio. The pitch is compelling: mTLS between every service with zero code changes, canary deploys at the traffic layer, automatic retries and circuit breakers, distributed tracing without instrumenting your application. All of this sounds like things you\u0026rsquo;d want. The question is whether you need a service mesh to get them, or whether the mesh introduces more complexity than it removes.\nThis post covers what a service mesh actually does, the two serious options (Istio and Linkerd), and the decision framework for whether you\u0026rsquo;ve reached the point where a mesh earns its cost.\nWhat a service mesh is, mechanically A service mesh adds a sidecar proxy (usually Envoy) next to every service instance. Every network call from your service goes through the proxy first. Every inbound call hits the proxy before it reaches your service. Your code doesn\u0026rsquo;t know the proxy exists: it makes normal HTTP or gRPC calls, and the proxy handles everything else.\nBefore mesh: Service A ──── HTTP ────→ Service B After mesh: Service A ──→ Sidecar A ──→ Sidecar B ──→ Service B (mTLS, retry, (mTLS, trace circuit break) context) The sidecars collectively form the \u0026ldquo;mesh\u0026rdquo;: they\u0026rsquo;re managed by a control plane (Istiod for Istio, the Linkerd control plane for Linkerd) that distributes configuration, certificates, and routing rules to every proxy.\nThe value proposition: all of the network-layer concerns (encryption, retries, timeouts, traffic splitting, observability) live in the proxy, not in your application code. You get them without changing a line of Go.\nWhat a service mesh actually gives you Mutual TLS (mTLS) without code changes Without a mesh, service-to-service traffic is usually unencrypted inside your cluster. Adding mTLS manually means each service manages its own certificates, rotates them, and verifies the other side\u0026rsquo;s certificate. That\u0026rsquo;s a significant operational burden: certificate management is the kind of thing that works until it doesn\u0026rsquo;t, and then every service starts rejecting connections because a certificate expired.\nA mesh handles this automatically. The control plane issues short-lived certificates to every sidecar, rotates them before expiry, and verifies identity on every connection. Your services get encrypted, authenticated traffic with zero code changes.\nTraffic management The mesh can split traffic by percentage: send 5% of requests to the new version of a service, 95% to the old. This is the canary deploy pattern implemented at the infrastructure layer instead of in your application or load balancer.\nIt can also handle retries, timeouts, and circuit breakers at the proxy level. If Service B is slow, Sidecar A retries the request without your code knowing. If Service B is down, Sidecar A opens a circuit and returns a fallback response immediately.\nObservability Every request that flows through the mesh is automatically traced. The sidecars inject trace headers (if you\u0026rsquo;re using OpenTelemetry or Jaeger), record request duration, and report success/failure rates. You get a dashboard of service-to-service latency and error rates without instrumenting your application.\nThis is the same value as the debugging post\u0026rsquo;s correlation IDs and tracing, but implemented in infrastructure instead of application code. The trade-off: the mesh gives you network-layer metrics (request duration, status codes, retry counts) but not application-layer metrics (business logic errors, database query latency, cache hit rates). You still need application-level observability.\nIstio vs. Linkerd Istio Istio is the most widely adopted service mesh. It uses Envoy as the sidecar, has a large feature set (traffic management, security, observability), and a large community.\nStrengths:\nFeature-complete. Traffic splitting, fault injection, rate limiting, policy enforcement; if the feature exists in the mesh space, Istio has it. Extensible via WebAssembly (WASM) filters. You can add custom logic to the proxy without forking it. Large ecosystem. Most Kubernetes tools and platforms integrate with Istio out of the box. Weaknesses:\nResource-heavy. The Istio control plane (Istiod) and the Envoy sidecars consume significant CPU and memory. On a small cluster, Istio\u0026rsquo;s overhead is a meaningful percentage of your total resources. Complex to operate. Certificate management, upgrade procedures, and debugging proxy issues require dedicated knowledge. Istio upgrades occasionally break things, and the release cadence is fast. Configuration is verbose. Istio\u0026rsquo;s CRDs (VirtualService, DestinationRule, Gateway, etc.) are powerful but numerous. Getting traffic routing right requires understanding several interacting resources. Linkerd Linkerd is a lighter-weight alternative. It uses a purpose-built proxy (not Envoy) called linkerd2-proxy, written in Rust.\nStrengths:\nLightweight. The proxy uses ~10MB of RAM and minimal CPU. The control plane is simpler and smaller than Istio\u0026rsquo;s. Simpler to operate. Fewer CRDs, simpler upgrade process, less configuration surface area. You can get mTLS and basic traffic management running in minutes, not hours. Strong defaults. Linkerd makes opinionated choices (automatic retries, automatic mTLS, automatic telemetry) that work out of the box without tuning. Weaknesses:\nSmaller feature set. No WASM extensibility, fewer traffic management options, no fault injection or rate limiting at the mesh layer. Smaller community. Fewer integrations, fewer blog posts, fewer Stack Overflow answers when you hit issues. Less customizable. If you need fine-grained control over proxy behavior, Linkerd\u0026rsquo;s simpler model might be too constraining. The honest comparison For most teams, the choice comes down to: Istio if you need the feature set, Linkerd if you need simplicity. If you\u0026rsquo;re adding a mesh primarily for mTLS and basic traffic management, Linkerd is the easier path. If you need rate limiting, fault injection, WASM extensibility, or complex traffic routing, Istio is the one with those features.\nBoth are production-ready. Both are used at scale by large companies. The risk of choosing either one is low; the bigger risk is adopting a mesh at all when you don\u0026rsquo;t need one.\nWhen you don\u0026rsquo;t need a mesh A service mesh is the answer to problems that are specifically about the network layer between services. If your problems aren\u0026rsquo;t network-layer problems, a mesh won\u0026rsquo;t help:\n\u0026ldquo;Our services are hard to debug.\u0026rdquo; A mesh gives you request tracing and latency metrics. It doesn\u0026rsquo;t give you application-level debugging: correlation IDs, structured logging, and the debugging workflow from the debugging post. Start there.\n\u0026ldquo;We need canary deploys.\u0026rdquo; You can do canary deploys with your cloud load balancer (weighted target groups in ALB, traffic splitting in Cloud Run) without a mesh. The mesh gives you a more sophisticated version, but the load balancer version works for most teams.\n\u0026ldquo;We need mTLS.\u0026rdquo; You can add mTLS with cert-manager and an ingress controller, or use a zero-trust network like Tailscale. It\u0026rsquo;s more manual than a mesh, but it\u0026rsquo;s also less infrastructure.\n\u0026ldquo;We want automatic retries and circuit breakers.\u0026rdquo; Libraries like sony/gobreaker for Go or resilience4j for Java give you retries and circuit breakers in application code. It\u0026rsquo;s more code than a mesh, but it\u0026rsquo;s code you control and understand.\nWhen you actually need one The mesh earns its cost when:\nYou have 15+ services and the operational burden of managing mTLS, traffic rules, and observability across all of them manually exceeds the operational burden of running the mesh. Below 15 services, the manual approach is usually cheaper.\nYou have strict security requirements that mandate encrypted service-to-service traffic and you can\u0026rsquo;t manage certificates manually. This is common in regulated industries (finance, healthcare).\nYou need sophisticated traffic management: canary deploys with automatic rollback based on error rates, fault injection for chaos engineering, rate limiting at the infrastructure layer. These are features that are hard to build yourself and easy with a mesh.\nYou\u0026rsquo;re on a platform team that will own the mesh. The mesh needs someone to operate it: upgrades, debugging proxy issues, writing traffic policies. If nobody owns it, it becomes a source of mysterious failures that nobody understands.\nThe best time to adopt a mesh is when you have a platform team that can own it and a scale that justifies the overhead. The worst time is when someone reads a blog post about mTLS and proposes Istio in standup without considering the operational cost.\nThe alternative: per-service libraries Before a mesh, many teams solve the same problems with libraries. A resilience library in your service handles retries, circuit breakers, and timeouts. A tracing library handles distributed tracing. A TLS library handles mTLS.\nThe trade-off: libraries require code changes in every service, and they lock you into a language. If your services are all in Go, a Go resilience library works. If someone writes a service in Python, you need a Python version too. A mesh is language-agnostic: the proxy handles everything regardless of what language your service is written in.\nFor polyglot teams, the mesh wins on consistency. For monolingual teams, libraries are simpler and have no infrastructure overhead.\nA service mesh is powerful infrastructure that solves real problems at scale. It\u0026rsquo;s also one of the most commonly adopted \u0026ldquo;because it exists\u0026rdquo; solutions in the microservices space. Start with application-level observability, library-based resilience, and your cloud\u0026rsquo;s built-in traffic management. When those stop being enough (when you have enough services that the manual approach is genuinely too much work), a mesh is the right next step. But \u0026ldquo;when those stop being enough\u0026rdquo; is usually later than you think.\n","permalink":"https://hanhpham.vercel.app/posts/microservices-service-mesh-when-you-actually-need-one/","summary":"A service mesh gives you mTLS, traffic management, and observability without changing application code. It also gives you another piece of infrastructure to operate, debug, and explain to new hires. Here\u0026rsquo;s how to tell whether the trade-off is worth it.","title":"Service Mesh: When You Actually Need One (and When You Don't)"},{"content":"\u0026ldquo;The API is slow\u0026rdquo; is the most common performance complaint and the least useful. Slow where? Slow in the network? In the database? In a loop that quadratic in the number of orders? Without a profiling tool, you\u0026rsquo;re guessing: moving things around, adding caches, rewriting code that might not be the bottleneck, and hoping you get lucky.\nGo ships with a profiling tool built in: pprof. It\u0026rsquo;s in the standard library, it works on any Go program, and it tells you exactly where your program spends its resources. This post covers the three profiles that matter most (CPU, memory, goroutine), how to collect them, and how to read the output without a PhD in computer science.\nThe three profiles you\u0026rsquo;ll actually use pprof can collect several types of profiles. Three of them cover 95% of performance investigations:\nCPU profile: where your program spends processor time. Use this when the symptom is \u0026ldquo;the service is slow\u0026rdquo; and you need to know which function is burning cycles.\nMemory (heap) profile: where your program allocates memory. Use this when the symptom is \u0026ldquo;the service uses too much RAM\u0026rdquo; or \u0026ldquo;GC is running too often\u0026rdquo; or \u0026ldquo;we\u0026rsquo;re getting OOM-killed.\u0026rdquo;\nGoroutine profile: where your goroutines are and what they\u0026rsquo;re doing. Use this when the symptom is \u0026ldquo;goroutine count is growing\u0026rdquo; or \u0026ldquo;the service is unresponsive\u0026rdquo; or \u0026ldquo;goroutine leak.\u0026rdquo;\nCollecting a profile: two ways Way 1: the net/http/pprof endpoint (for running services) The simplest way to profile a running Go service: import net/http/pprof and expose it on an HTTP endpoint. If your service already runs an HTTP server, this is a two-line change:\nimport _ \u0026#34;net/http/pprof\u0026#34; // in your main or init go func() { http.ListenAndServe(\u0026#34;localhost:6060\u0026#34;, nil) }() Now your service exposes profiling data at localhost:6060/debug/pprof/. Collect a 30-second CPU profile:\ngo tool pprof http://localhost:6060/debug/pprof/profile?seconds=30 This downloads the profile and opens an interactive shell. The service runs normally while the profile is being collected: no restart, no code change, no special build flags.\nSecurity note: never expose the pprof endpoint to the public internet. It gives anyone who can reach it deep insight into your program\u0026rsquo;s behavior. Bind it to localhost, or put it behind an auth proxy, or only enable it in non-production environments.\nWay 2: runtime/pprof in code (for benchmarks and CLIs) For programs that don\u0026rsquo;t run an HTTP server (CLI tools, batch jobs, tests), write the profile directly:\nimport \u0026#34;runtime/pprof\u0026#34; func main() { f, _ := os.Create(\u0026#34;cpu.prof\u0026#34;) pprof.StartCPUProfile(f) defer pprof.StopCPUProfile() // ... your code here } Then analyze with:\ngo tool pprof cpu.prof Reading a CPU profile You\u0026rsquo;ve collected a profile. You\u0026rsquo;re staring at a wall of text. Here\u0026rsquo;s what to look for.\nThe interactive pprof shell shows you functions sorted by how much CPU time they consume. The most useful commands:\n(pprof) top 20 Showing nodes accounting for 3.2s, 80% of 4s flat flat% sum% cum cum% 1.2s 30.0% 30.0% 1.2s 30.0% runtime.mallocgc 0.8s 20.0% 50.0% 0.8s 20.0% runtime.memclrNoHeapPointers 0.4s 10.0% 60.0% 0.6s 15.0% encoding/json.Marshal ... flat is the time spent in the function itself (not its callees). cum is the time spent in the function plus everything it calls. A function with high flat is doing expensive work directly. A function with high cum but low flat is calling expensive functions; the bottleneck is in one of its callees.\nThe pattern to look for: high cum with low flat. That means the function is slow because something it calls is slow. Drill down with list \u0026lt;function\u0026gt; to see which line:\n(pprof) list json.Marshal Showing nodes accounting for 0.6s, 15% of 4s Total: 4s ROUTINE ======================== in encoding/json 0.4s 0.4s (flat, cum) 10.0% of Total . . 87: e.init() . . 88: e.p = e.pretty 0.4s 0.4s 89: e.reflectValue(v, encOpts{escapeHTML: true}) Line 89, reflectValue, is where the time goes. You\u0026rsquo;re spending 10% of total CPU time in JSON reflection. The fix: reduce the amount of data you\u0026rsquo;re marshaling, or use a faster JSON library like json-iterator or sonic.\nReading a memory profile A heap profile shows where your program allocates memory. The interactive commands are the same (top, list, web) but the numbers mean something different:\n(pprof) top 10 Showing nodes accounting for 128MB, 64% of 200MB flat flat% sum% cum cum% 64MB 32.0% 32.0% 64MB 32.0% github.com/company/service/models.(*Order).ToJSON 32MB 16.0% 48.0% 32MB 16.0% database/sql.(*DB).Query 16MB 8.0% 56.0% 16MB 8.0% fmt.Sprintf flat here is bytes allocated, not time. The function at the top (ToJSON) allocates 64MB, half your total heap. Drill down:\n(pprof) list ToJSON 64MB 64MB 42: func (o *Order) ToJSON() []byte { . . 43: data, _ := json.Marshal(o) . . 44: return data Every call to ToJSON allocates a new byte slice via json.Marshal. If you\u0026rsquo;re marshaling thousands of orders per request, that\u0026rsquo;s a lot of short-lived allocations that the garbage collector has to clean up. The fix: reuse buffers with sync.Pool, or stream the JSON instead of materializing the whole thing in memory.\nIn-use vs. alloc The heap profile has two views: in-use (what\u0026rsquo;s allocated right now) and alloc (what was allocated since the program started). Switch between them:\n(pprof) alloc_space # total allocated since start (pprof) inuse_space # what\u0026#39;s live right now A function that shows up in alloc but not in-use allocates a lot of memory that gets garbage collected quickly. That\u0026rsquo;s a GC pressure problem, not a memory leak. A function that shows up in in-use allocates memory that never gets freed. That\u0026rsquo;s a leak.\nReading a goroutine profile A goroutine profile shows every live goroutine and where it\u0026rsquo;s blocked:\n(pprof) top 10 1008 @ 0x43e236 0x40a3c5 0x40a3c5 0x46f8a5 0x4712b8 0x6d3f1a 0x472160 # 0x6d3f1a database/sql.(*DB).Query+0x13a /usr/local/go/src/sql/sql.go:1924 # 0x472160 service.(*OrderService).GetOrders+0x40 /app/service/orders.go:87 If you see hundreds of goroutines all stuck in the same place (say, sql.(*DB).Query), you have a database connection pool exhaustion problem. Your goroutines are all waiting for a database connection that isn\u0026rsquo;t available. The fix: increase the pool size, or find the query that\u0026rsquo;s holding connections open too long.\nIf goroutine count is growing over time (collect two profiles a minute apart and compare), you have a goroutine leak: goroutines that are spawned but never finish. The goroutine profile shows where they\u0026rsquo;re stuck, which tells you why they\u0026rsquo;re not finishing.\nThe practical workflow Here\u0026rsquo;s the actual process when your service is slow:\nCollect a CPU profile while the service is slow. Hit the pprof endpoint, wait 30 seconds, download the profile.\nRun top 20 to see the top consumers. Look for functions with high cum: those are the bottlenecks.\nDrill down with list to see which line in the function is expensive. Is it a database call? A JSON marshal? A regex? A lock?\nFix the specific thing that\u0026rsquo;s slow. Not \u0026ldquo;rewrite the service.\u0026rdquo; Not \u0026ldquo;add a cache.\u0026rdquo; The specific function, the specific line.\nCollect another profile after the fix to verify it actually helped.\nThe most common mistake: optimizing the wrong thing. You see \u0026ldquo;30% in runtime.mallocgc\u0026rdquo; and think \u0026ldquo;GC is slow.\u0026rdquo; But mallocgc is high because you\u0026rsquo;re allocating too much; the fix isn\u0026rsquo;t faster allocation, it\u0026rsquo;s fewer allocations. Find what\u0026rsquo;s allocating (the alloc_space profile), not what\u0026rsquo;s collecting.\nThe second most common mistake: profiling in development instead of production. A CPU profile under no load tells you where your program spends time when nothing is happening. Profile under real traffic to find the real bottlenecks. The pprof endpoint on localhost is safe for production; just don\u0026rsquo;t expose it publicly.\nA worked example: finding a quadratic loop MusicCorp\u0026rsquo;s order summary endpoint takes 2 seconds for 1,000 orders and 20 seconds for 10,000 orders. That\u0026rsquo;s quadratic: doubling the input quadruples the time.\nCPU profile:\n(pprof) top 5 8.0s 40.0% 40.0% 8.0s 40.0% service.(*OrderSummary).buildIndex 4.0s 20.0% 60.0% 4.0s 20.0% runtime.makeslice Drill down:\n(pprof) list buildIndex 8.0s 8.0s 31: for i, order := range orders { 0.0s 0.0s 32: for j, other := range orders { 8.0s 0.0s 33: if order.CustomerID == other.CustomerID { 0.0s 0.0s 34: index[order.CustomerID] = append(...) 0.0s 0.0s 35: } 0.0s 0.0s 36: } Lines 32-36: a nested loop comparing every order to every other order. That\u0026rsquo;s O(n²). The fix: build the index with a single pass using a map:\nfor _, order := range orders { index[order.CustomerID] = append(index[order.CustomerID], order) } Single pass, O(n). 20 seconds becomes 200 milliseconds. The profile told you exactly where the problem was. No guessing required.\npprof is the tool that turns \u0026ldquo;the service is slow\u0026rdquo; into \u0026ldquo;line 33 in buildIndex is an O(n²) nested loop.\u0026rdquo; It\u0026rsquo;s built into Go, it works on running services with zero code changes (via the HTTP endpoint), and it takes thirty seconds to collect. The next time something is slow, reach for the profile before reaching for the keyboard.\nIf you\u0026rsquo;re interested in seeing how profiling works at a lower level (reading binary output, tracing execution without source code), check out the reverse engineering series, which covers Ghidra, GDB, and binary analysis.\n","permalink":"https://hanhpham.vercel.app/posts/go-performance-profiling-with-pprof/","summary":"Your Go service is slow, but you don\u0026rsquo;t know where. Guessing is not a strategy. pprof tells you exactly where your program spends its time, memory, and CPU; here\u0026rsquo;s how to use it.","title":"Go Performance Profiling with pprof: Finding the Slow Part"},{"content":"In the previous post we drew boundaries around MusicCorp\u0026rsquo;s services using coupling and cohesion. Now those services need to actually talk. And here\u0026rsquo;s where most teams make their first real mistake: they open a discussion about technology (REST vs. gRPC vs. Kafka) before they\u0026rsquo;ve settled a much more basic question. Two questions, really:\nDoes the caller block waiting for this to finish, or carry on? Is this a request aimed at a specific service, or a broadcast that nobody in particular has to be listening for? Those two axes give you four communication styles, plus a fifth (sharing data through a common store) that\u0026rsquo;s easy to miss because it barely looks like communication at all. Get this decision right first, and the technology choice in the next post becomes a much shorter conversation.\nWhy \u0026ldquo;just make it a network call\u0026rdquo; doesn\u0026rsquo;t work It\u0026rsquo;s tempting to treat a call to another service like a method call on an object: cross the process boundary, get an answer back, move on. Three things break that illusion immediately:\nPerformance. An in-process call can be inlined away by the compiler. An inter-process call means serializing data, sending packets, and waiting, milliseconds where a local call was nanoseconds. A design that makes sense as 1,000 in-process calls is a bad idea as 1,000 network calls. Interface changes stop being atomic. Change a method signature in-process and your IDE fixes every call site in the same commit. Change a service\u0026rsquo;s interface and the caller is a separately-deployed process that finds out on its own schedule. Failure gets non-deterministic. A local call either works or throws. A network call can time out, arrive twice, arrive out of order, or get a response that never makes it back because the caller died in the meantime. Distributed systems research breaks this down into crash, omission, timing, response, and (worst of all) arbitrary failures, where the parties involved can\u0026rsquo;t even agree that something went wrong. None of that is a reason to avoid inter-process communication; it\u0026rsquo;s a reason to design for it deliberately, which is what the rest of this post is about.\nSynchronous blocking: familiar, and dangerous in chains The simplest mental model: OrderProcessor calls Loyalty to add points, and blocks until it hears back. This is how most of us learned to program (one line waits for the previous one), so it\u0026rsquo;s the natural first choice when moving off a monolith.\nThe cost is temporal coupling: both the caller and callee, and specifically these instances of them, have to be up at the same moment. If Loyalty is slow, OrderProcessor is slow. If Loyalty is down, the call fails and OrderProcessor has to decide what to do about it right now.\nThis gets genuinely dangerous once calls start chaining. Picture a fraud check on checkout:\nOrderProcessor --\u0026gt; Payment --\u0026gt; FraudDetection --\u0026gt; Customer If all four hops are synchronous and blocking, a hiccup anywhere in that chain fails the whole operation, and every hop in between is holding a connection open the entire time, a good way to run out of available connections under load. Two fixes are worth knowing before you reach for asynchronous communication as the default cure: shorten the chain (does FraudDetection really need to be in the critical path, or could it run in the background and flag problem accounts ahead of time?), or replace the blocking calls with a nonblocking style, which is the next stop.\nAsynchronous nonblocking: decoupled, but a different way of thinking With async communication, the caller fires off the call and keeps working without waiting for a response. This buys you temporal decoupling (the receiving service doesn\u0026rsquo;t need to be reachable at the exact moment the call is made), and it\u0026rsquo;s close to mandatory for anything long-running. Packaging and dispatching an order might take hours or days; you cannot hold a synchronous connection open for that.\nThe cost is complexity you don\u0026rsquo;t get with a blocking call: if the response comes back later, does it come back to the same instance that made the request? What if that instance is gone by then? You typically need to persist enough state that whichever instance picks up the response can reconstruct what it was for.\nA word of caution before you get excited about async everywhere: it trades one set of headaches (blocking, cascading failure) for another (out-of-order delivery, duplicate messages, \u0026ldquo;did that response go anywhere?\u0026rdquo;). It\u0026rsquo;s not simpler, just differently complex. Good monitoring and a correlation ID on every message are not optional extras here: they\u0026rsquo;re how you\u0026rsquo;ll debug the \u0026ldquo;where did this message go\u0026rdquo; question at 2 a.m.\nRequest-response: the caller wants a specific answer Orthogonal to sync/async is a second axis: is the caller asking a specific service to do something and expecting to hear back? That\u0026rsquo;s request-response, and it works in either flavor:\nSynchronous request-response: Chart asks Inventory for current stock levels over HTTP, blocks, gets an answer. Asynchronous request-response: OrderProcessor puts a \u0026ldquo;reserve stock\u0026rdquo; message on a queue; Inventory picks it up whenever it\u0026rsquo;s free, does the work, and puts the response on a reply queue that OrderProcessor reads from. Request-response is the right shape whenever you genuinely need the result before you can continue, or you need to know whether something failed so you can retry or compensate. If either of those is true, request-response fits; sync vs. async is then a question of whether you can afford to block.\nOne practical trap worth flagging explicitly: if you need results from several independent request-response calls before proceeding (say, checking price from three different stockists), running them in sequence costs you the sum of their latencies. Running them in parallel costs you only the slowest one. This sounds obvious written down, but it\u0026rsquo;s an easy thing to get wrong by default when the code just calls three things in a row because that\u0026rsquo;s how it was written.\nEvent-driven: the inversion that takes getting used to This is the odd one out, and it\u0026rsquo;s worth sitting with because the mental model is genuinely inverted from request-response. Instead of asking a specific service to do something, a service just broadcasts a fact: \u0026ldquo;this happened.\u0026rdquo; Warehouse fires an event when a package is packed. It does not know or care who\u0026rsquo;s listening: Notifications might send an email, Inventory might adjust stock counts, both, or neither. The emitter is unaware of, and doesn\u0026rsquo;t need to be aware of, who consumes its events.\nThat inversion is exactly what makes event-driven collaboration so loosely coupled: with request-response, the caller has to know what the downstream service can do, a form of domain coupling. With events, the emitter knows nothing about its consumers, so there\u0026rsquo;s nothing to couple to.\nThe catch is what goes inside the event. Two options:\nJust an ID: consumers that need more than the ID have to call back to fetch it, which reintroduces domain coupling and can hammer the source service if many consumers all react to the same event. Fully detailed: put in everything a consumer would reasonably need, the same as you would for a request-response payload. This is generally the better default, at the cost of the event becoming a wider contract you now have to maintain (remove a field later, and you might break someone quietly depending on it). Events are also, by their nature, always asynchronous: there\u0026rsquo;s no such thing as a synchronous broadcast, since the emitter by design doesn\u0026rsquo;t wait for anyone.\nThe pattern you don\u0026rsquo;t notice: communication through common data The fifth style barely feels like \u0026ldquo;communication\u0026rdquo; because it\u0026rsquo;s so indirect: one service drops data somewhere (a file, a data lake, a shared table) and one or more other services pick it up later, usually by polling. Despite feeling informal, this is arguably the most common integration pattern in existence, especially for large data volumes or when you need interoperability with something that can\u0026rsquo;t speak your API\u0026rsquo;s protocol (an old mainframe can usually still read a file, even if it\u0026rsquo;s never heard of gRPC).\nThe failure mode to watch is the same common coupling from the last post: if multiple services both read and write the shared store, you\u0026rsquo;ve built exactly the tangled shared-database problem coupling analysis was supposed to help you avoid. Keep the flow of information one-directional (publisher writes, consumers only read) and this pattern is genuinely useful rather than an accident waiting to surface.\nMix and match, on purpose A real microservice architecture is not \u0026ldquo;we do REST\u0026rdquo; or \u0026ldquo;we do events\u0026rdquo;; it\u0026rsquo;s a deliberate mix. It\u0026rsquo;s entirely normal for one service to expose a synchronous request-response API for placing an order and fire events when the order\u0026rsquo;s state changes, serving both an immediate caller and any number of interested listeners. The skill isn\u0026rsquo;t picking one style company-wide; it\u0026rsquo;s recognizing, interaction by interaction, which of these shapes actually fits what you\u0026rsquo;re building.\nThese are technology-agnostic patterns: none of them mandate REST, gRPC, or Kafka specifically. Next up: which actual technology fits which pattern, and where the popular choices (REST, gRPC, GraphQL, message brokers) genuinely differ.\n","permalink":"https://hanhpham.vercel.app/posts/microservices-five-ways-to-talk/","summary":"Before picking REST, gRPC, or Kafka, you need to answer a more basic question: does this interaction block, and does the caller expect an answer? Get that wrong and no technology choice saves you.","title":"The Five Ways Microservices Talk to Each Other"},{"content":"This is part two of a series on microservices architecture. If you\u0026rsquo;re not sure whether microservices are the right call for your team yet, start with the decision checklist before reading further.\nSay you\u0026rsquo;re building the backend for an online CD retailer. We\u0026rsquo;ll call it MusicCorp, and we\u0026rsquo;ll keep coming back to it in this series. Orders need payment taken, stock reserved, packages shipped, loyalty points awarded. Somebody on the team draws a box around each of those verbs, calls it a microservice, and ships it. Six months later, changing how loyalty points work requires touching four different services in lockstep. What went wrong?\nAlmost always, it\u0026rsquo;s the boundary, not the technology. Kubernetes, gRPC, and a service mesh won\u0026rsquo;t save you from a bad cut. This post is about the three ideas that actually predict whether a boundary will hold up: information hiding, cohesion, and coupling, plus the specific kinds of coupling worth telling apart, because \u0026ldquo;avoid coupling\u0026rdquo; is useless advice until you know which coupling you\u0026rsquo;re looking at.\nInformation hiding: the one rule underneath everything else David Parnas, writing about module design decades before \u0026ldquo;microservice\u0026rdquo; was a word, put it simply:\nThe connections between modules are the assumptions which the modules make about each other.\nEvery assumption one service makes about another\u0026rsquo;s internals is a thread that will eventually snap when someone changes those internals. The fewer assumptions, the more freely each side can change. A microservice\u0026rsquo;s whole value proposition (deploy this one thing without deploying anything else) depends on ruthlessly hiding as much as possible behind its interface: internal data structures, database schema, business logic, all of it. Expose only the minimum needed to satisfy consumers.\nCohesion: the code that changes together, stays together Cohesion asks a different question than coupling: it\u0026rsquo;s about what\u0026rsquo;s inside your boundary, not what crosses it. The pithiest definition: the code that changes together should live together. If a single business change (say, \u0026ldquo;loyalty points now expire after 12 months\u0026rdquo;) requires edits spread across three services, that\u0026rsquo;s weak cohesion; the related behavior never had a coherent home. Strong cohesion means one change, one deploy.\nCoupling and cohesion aren\u0026rsquo;t independent: they\u0026rsquo;re two views of the same underlying question, just measured from inside vs. outside the boundary. Larry Constantine\u0026rsquo;s law says it cleanly: a structure is stable if cohesion is strong and coupling is low. Neither one alone is sufficient.\nCoupling comes in flavors, and they are not equally bad This is the part most teams skip, and it\u0026rsquo;s the part that actually matters. \u0026ldquo;Loose coupling good, tight coupling bad\u0026rdquo; doesn\u0026rsquo;t tell you what to do when you\u0026rsquo;re staring at a design and trying to decide if it\u0026rsquo;s fine. Here are four kinds of coupling you\u0026rsquo;ll run into between microservices, ordered from least to most dangerous.\nDomain coupling: usually fine, in moderation Domain coupling is simply one service calling another because it needs that other service\u0026rsquo;s functionality. OrderProcessor calls Warehouse to reserve stock and Payment to take money. This is largely unavoidable (a system made of collaborating services has to collaborate) and it\u0026rsquo;s considered the loosest, most acceptable form of coupling.\nThe warning sign isn\u0026rsquo;t the coupling itself, it\u0026rsquo;s the shape of it: if one service depends on a long list of downstream services, that\u0026rsquo;s often a symptom that too much logic and responsibility has piled up in the caller. Keep the fan-out small, and keep what you send across the boundary to the minimum the callee actually needs; information hiding again.\nPass-through coupling: the sneaky one This happens when a service passes data through to a second caller purely because a third, further-downstream service needs it. Picture OrderProcessor sending a ShippingManifest to Warehouse, which does nothing with it except forward it to Shipping. Now OrderProcessor has to know about a data shape that belongs, conceptually, to a service two hops away. Change what Shipping needs, and the change potentially ripples all the way back to OrderProcessor; three services now need a coordinated release for what should have been an internal detail of one.\nThe fix is usually one of:\nBypass the intermediary: have the caller talk to the real owner directly. Trades pass-through coupling for a bit more domain coupling, which is a good trade, but only if it doesn\u0026rsquo;t push logic that belonged to the intermediary back up into the caller. Let the intermediary own the shape: have Warehouse collect what it needs and construct the ShippingManifest itself, so Shipping\u0026rsquo;s contract changes become invisible to OrderProcessor. Treat the payload as an opaque blob: OrderProcessor still sends the manifest through Warehouse, but Warehouse never looks inside it, just relays it. This doesn\u0026rsquo;t eliminate the coupling between the two endpoints, but it does mean the middle service never needs to change when the shape does. Common coupling: fine for read-only reference data, risky otherwise Common coupling is what you get when two or more services read and write the same shared data (classically, a shared database table). Multiple services reading static, rarely-changing reference data (country codes, tax rates) from one store is relatively benign, because that data barely changes and nobody\u0026rsquo;s fighting over write access.\nIt gets dangerous the moment multiple services both write to the same structure. Imagine OrderProcessor and Warehouse both updating a Status column on the same Order row: one setting PLACED/PAID/COMPLETED, the other setting PICKING/SHIPPED. Nothing stops an invalid transition like PLACED → SHIPPED from slipping through, because neither service has a complete view of what\u0026rsquo;s allowed. The fix is to give the state machine a single owner: an Order service that both callers request changes from, and that can reject a request that violates its own rules. For the practical patterns of how to structure data when services can\u0026rsquo;t share a database (API composition, CQRS, event sourcing), see database per service.\nIf a \u0026ldquo;service\u0026rdquo; is really just a thin wrapper over database CRUD (every request maps straight to an update, no rules applied), that\u0026rsquo;s a sign the logic that should live there has leaked out into every caller instead. You\u0026rsquo;ve traded one strongly-coupled shared table for several services all independently guessing what\u0026rsquo;s a valid state transition.\nContent coupling: just don\u0026rsquo;t Content coupling is common coupling\u0026rsquo;s uglier sibling: instead of a known shared dependency, an outside service reaches directly into another service\u0026rsquo;s internal storage and mutates it, bypassing the owning service\u0026rsquo;s API entirely. Say Warehouse writes directly to the Order table instead of calling the Order service.\nThe difference from common coupling is subtle but important: with common coupling, everyone at least knows they share an external dependency they don\u0026rsquo;t fully control. With content coupling, the lines of ownership disappear. The Order table becomes part of an external contract nobody agreed to, information hiding is gone, and you\u0026rsquo;re now trusting that Warehouse\u0026rsquo;s idea of \u0026ldquo;valid state transition\u0026rdquo; exactly matches the Order service\u0026rsquo;s, with zero way to enforce it. This is sometimes called pathological coupling for a reason. Avoid it outright.\nA quick word on temporal coupling One more form worth knowing, because it comes up constantly once you start picking communication styles (the subject of the next post): temporal coupling is when two services both need to be up and reachable at the same instant for an operation to succeed. A synchronous HTTP call from OrderProcessor to Warehouse is temporally coupled: if Warehouse is down, the whole operation fails right then. It\u0026rsquo;s not inherently bad, but as call chains grow, temporal coupling compounds into cascading failures. Asynchronous communication is one of the main tools for loosening it.\nNaming things the way your users do One more foundational habit worth adopting alongside coupling analysis: ubiquitous language, borrowed from Domain-Driven Design. Use the same terms in your code that domain experts use when they talk about the business. The alternative (a generic, one-size-fits-all data model where every real-world concept becomes a vague Arrangement or Entity) forces constant mental translation between what the business means and what the code says, and that translation tax gets paid by every developer, forever.\nA related DDD concept, the aggregate, is worth carrying forward too: an aggregate is a real domain concept with its own identity and lifecycle (an Order, an Invoice), modeled as a self-contained unit whose state transitions are managed together, in one place, by one service. An aggregate gets to say no to an invalid request. That idea (a service guarding its own state machine rather than letting callers dictate it directly) is exactly what breaks the common-coupling trap above, and it\u0026rsquo;s worth keeping in your back pocket for the rest of this series.\nBoundaries are the foundation, but they don\u0026rsquo;t tell you how two well-bounded services should actually talk to each other at runtime: synchronously, asynchronously, as a request or as a broadcast event. That\u0026rsquo;s next.\n","permalink":"https://hanhpham.vercel.app/posts/microservices-boundaries-coupling-cohesion/","summary":"Before you draw a single service boundary, you need a vocabulary for why some splits age well and others turn into a distributed monolith. Coupling and cohesion are that vocabulary.","title":"What Actually Makes a Good Microservice Boundary"},{"content":"Every six months, someone on the team reads a blog post about how Company X scaled to ten million users with microservices, and suddenly there\u0026rsquo;s a proposal on the whiteboard to split the monolith. The proposals always sound reasonable. The architecture diagrams always look clean. And about three months into the actual migration, someone quietly mutters in standup: \u0026ldquo;wait, why did we do this again?\u0026rdquo;\nThis post is the one I wish someone had handed me before I went through that cycle the first time. It won\u0026rsquo;t tell you how to build microservices (the rest of this series covers that); it\u0026rsquo;ll tell you whether you should be building them at all.\nThe uncomfortable truth: microservices solve an organizational problem Microservices aren\u0026rsquo;t a technology pattern. They\u0026rsquo;re an organizational pattern that happens to require technology to implement. The core value isn\u0026rsquo;t \u0026ldquo;services are smaller so they\u0026rsquo;re easier to understand\u0026rdquo;: a well-structured monolith with clean module boundaries gives you that. The core value is independent deployability: Team A can ship their changes on Tuesday without coordinating with Team B\u0026rsquo;s release on Thursday.\nThat only matters when you have enough teams that coordination is genuinely slowing you down. If your engineering team can fit around a single table, you almost certainly don\u0026rsquo;t have this problem. A well-structured monolith with clear internal boundaries will let you move faster, deploy more confidently, and spend your time on features instead of infrastructure.\nThe question isn\u0026rsquo;t \u0026ldquo;are microservices better?\u0026rdquo;; it\u0026rsquo;s \u0026ldquo;do I have the problem microservices solve?\u0026rdquo; If your deploy bottleneck isn\u0026rsquo;t team coordination, you\u0026rsquo;re about to trade a mild inconvenience for a significant amount of operational complexity.\nWhat you\u0026rsquo;re actually giving up When you split a monolith into services, here\u0026rsquo;s what disappears, and most blog posts about microservices either skip this section or bury it in a paragraph near the end.\nACID transactions across business operations. In a monolith, placing an order and reserving stock can be one database transaction. If the stock reservation fails, the order never happened. With microservices, you now have two separate transactions that can each succeed or fail independently. You can get most of the way back with patterns like sagas (covered later in this series), but \u0026ldquo;most of the way back\u0026rdquo; is not \u0026ldquo;all the way back.\u0026rdquo; You\u0026rsquo;re accepting eventual consistency where you once had strong consistency, and that trade-off shows up in surprising places, like the customer who sees their order confirmation before the payment fails.\nA single, readable stack trace. Debugging a monolith means reading a stack trace from top to bottom. Debugging microservices means grepping across three services\u0026rsquo; logs, correlating by request ID, and trying to reconstruct a timeline that spans multiple processes. At 2 AM. With incomplete logs because one service was deployed with a logging level that\u0026rsquo;s too quiet. Distributed tracing tools like OpenTelemetry help, but they\u0026rsquo;re additional infrastructure you now have to run, configure, and teach your team to use. If you want to see what this debugging experience actually looks like (and what tooling makes it survivable), see debugging microservices.\nOne deploy target. \u0026ldquo;Ship it\u0026rdquo; in a monolith means pushing one thing. In a microservices world, \u0026ldquo;ship it\u0026rdquo; means deploying one service and praying that every other service it talks to can handle the new contract. Versioning, backward compatibility, and coordinated rollouts become your problem. If Service A sends a field that Service B used to ignore but now treats as required, you\u0026rsquo;ve just broken production in a way that won\u0026rsquo;t show up until the right combination of requests hits the new code path. If you want the playbook for making independent deploys actually work: rolling, blue-green, canary, feature flags, and the CI/CD structure that ties them together. See deploying without the fear.\nSimple onboarding. A new developer can read an entire monolith\u0026rsquo;s codebase in a week. A new developer looking at a microservice architecture sees an opaque web of services, each with its own repository, deployment pipeline, and implicit contracts with neighbors they haven\u0026rsquo;t met yet. \u0026ldquo;How does an order get processed?\u0026rdquo; used to be answered by reading one file. Now it\u0026rsquo;s a guided tour across four services and an event bus.\nCheap local development. Running a monolith locally means go run . and you\u0026rsquo;re done. Running a microservice architecture locally means Docker Compose with fifteen containers, or a shared staging environment that\u0026rsquo;s always half-broken because someone else\u0026rsquo;s changes broke the inter-service wiring. You can paper over this with a good developer platform, but that platform is itself a significant engineering investment. If you want the practical playbook for making this work (Docker Compose profiles, service virtualization, the layered testing strategy), see local dev without Docker Compose hell.\nThe checklist: six questions before you split If you can answer \u0026ldquo;yes\u0026rdquo; to most of these, microservices are probably worth evaluating seriously. If most answers are \u0026ldquo;no,\u0026rdquo; a well-structured monolith will serve you better.\n1. Do you have more than three teams that need to deploy independently? This is the big one. Microservices exist so that Team A can ship without waiting for Team B. If you have one team, or even three teams that are closely coordinated, the coordination overhead of independent services exceeds the coordination overhead of a shared codebase. You\u0026rsquo;ll spend more time managing service boundaries than you ever saved by decoupling deploys.\n2. Is your monolith\u0026rsquo;s deploy cycle measured in weeks, not days? If your team can deploy the monolith once a day and the process takes fifteen minutes, your deploy process isn\u0026rsquo;t the bottleneck; your features are. Splitting a fast-deploying monolith into microservices doesn\u0026rsquo;t make you faster; it makes you slower, because now every feature touches more deploy pipelines and more integration surfaces.\n3. Have you tried splitting the monolith\u0026rsquo;s modules first? Before reaching for network boundaries, try code boundaries. Extract a well-defined module into its own package with a clean internal API. If you can get the benefits of separation (independent understanding, clear ownership, testable in isolation) without the network, you should. The vast majority of \u0026ldquo;we need microservices\u0026rdquo; problems are actually \u0026ldquo;we need better module boundaries\u0026rdquo; problems.\n4. Do different parts of the system have genuinely different scaling requirements?\nIf your image processing pipeline needs ten times the compute of your user management service, and those requirements will continue to diverge, running them as separate services with independent scaling makes sense. If everything scales together, independent services don\u0026rsquo;t buy you anything; you\u0026rsquo;re just running the same infrastructure with extra network calls in between.\n5. Do you have the operational maturity to handle distributed failures? Microservices fail differently than monoliths. Network calls time out. Messages get delivered twice. Services go down independently while the rest of the system keeps running. Before you adopt microservices, you need:\nStructured, centralized logging (not just log.Printf to stdout) Distributed tracing across service boundaries Health checks and circuit breakers on every inter-service call An on-call rotation that can handle \u0026ldquo;the system is 80% up\u0026rdquo; instead of \u0026ldquo;the system is down\u0026rdquo; If you don\u0026rsquo;t have these yet, building them as part of a microservices migration is like learning to swim by jumping into the ocean. Build the observability stack first; if you\u0026rsquo;re not sure where to start, see debugging microservices for the minimum viable setup. If you can\u0026rsquo;t justify the observability investment on its own, you probably don\u0026rsquo;t need microservices either.\n6. Are you solving a real scaling problem, or a theoretical one? \u0026ldquo;We might need to scale this independently someday\u0026rdquo; is not a reason to take on the cost of microservices today. Scaling requirements are the easiest thing to defer. If you\u0026rsquo;re wrong about needing independent scaling, you\u0026rsquo;ve paid a permanent cost for a problem that never materialized. Wait until the scaling need is real and measurable, then split.\nWhat to do instead: the modular monolith If you answered \u0026ldquo;no\u0026rdquo; to most of the checklist above but still feel the pain of a tangled codebase, the answer is almost certainly a modular monolith: the same deployment unit, but with clear internal boundaries enforced at the code level.\nThe rules are simple:\nEach module gets its own package or directory, with a public API that other modules call. No reaching into another module\u0026rsquo;s internals. Modules communicate through their public API, not through shared database tables. If Module A needs data from Module B, it calls Module B\u0026rsquo;s function; it doesn\u0026rsquo;t query Module B\u0026rsquo;s table directly. Modules can be tested in isolation. If you can\u0026rsquo;t unit test a module without spinning up the whole application, the boundary isn\u0026rsquo;t clean enough. One team owns each module (or at least one module has a clear primary owner). Ownership means \u0026ldquo;responsible for the module\u0026rsquo;s API, internal design, and the data it manages.\u0026rdquo; A modular monolith gives you 80% of the organizational benefit of microservices (clear ownership, independent understanding, testability) without any of the operational cost. And critically, it leaves the door open to extract a module into a service later, when and if the real need arises. The module boundary becomes the future service boundary, pre-drawn.\nThe companies that successfully adopt microservices almost all went through a modular monolith phase first. They learned where the real boundaries were by running a monolith with enforced internal boundaries, then extracted the modules that genuinely needed independent deployment. The ones that jumped straight to microservices from a tangled monolith mostly ended up with a tangled distributed system: all the coupling they had before, plus network latency.\nIf you\u0026rsquo;re still reading: you might actually need them If you looked at the checklist and thought \u0026ldquo;yes, yes, yes, no, yes, yes\u0026rdquo;, if you genuinely have multiple teams stepping on each other\u0026rsquo;s deploys, if your monolith\u0026rsquo;s module boundaries have been tried and aren\u0026rsquo;t enough, if you have the operational stack to handle distributed failure modes, then the rest of this series is for you.\nWe start with the hardest and most consequential decision: how to draw the boundaries so they actually hold up. That\u0026rsquo;s where most microservice architectures succeed or fail, and it\u0026rsquo;s where we\u0026rsquo;ll begin: what actually makes a good microservice boundary.\n","permalink":"https://hanhpham.vercel.app/posts/should-you-even-use-microservices/","summary":"Microservices are a solution to organizational scaling problems, not a technology choice. Here\u0026rsquo;s how to tell whether you actually have the problem they solve, before you pay the cost.","title":"Should You Even Use Microservices?"},{"content":"I build things with code and sometimes break things on purpose.\nBy day, I\u0026rsquo;m a software engineer, working with Go, distributed systems, and the occasional smart contract. By night, I\u0026rsquo;m usually deep in a CTF challenge, reading manga, or watching an AI do something it shouldn\u0026rsquo;t be able to do.\nThis blog is where I write down the things I learn so I don\u0026rsquo;t forget them. Mostly technical: reverse engineering, microservices, security, whatever I\u0026rsquo;m currently obsessed with. Sometimes not.\nWhat I\u0026rsquo;m into AI software engineering blockchain hacking games manga I believe the best way to learn something is to build it, break it, and then write about it so future-me doesn\u0026rsquo;t have to figure it out again.\nOpen to contract, consulting, and engineering work in distributed systems, backend architecture, and security: reach me directly at hanhphamit@gmail.com.\nIf any of this sounds interesting, the posts are where the good stuff lives.\n","permalink":"https://hanhpham.vercel.app/about/","summary":"\u003cp\u003eI build things with code and sometimes break things on purpose.\u003c/p\u003e\n\u003cp\u003eBy day, I\u0026rsquo;m a software engineer, working with Go, distributed systems,\nand the occasional smart contract. By night, I\u0026rsquo;m usually deep in a CTF\nchallenge, reading manga, or watching an AI do something it shouldn\u0026rsquo;t be\nable to do.\u003c/p\u003e\n\u003cp\u003eThis blog is where I write down the things I learn so I don\u0026rsquo;t forget them.\nMostly technical: reverse engineering, microservices, security, whatever\nI\u0026rsquo;m currently obsessed with. Sometimes not.\u003c/p\u003e","title":"About"},{"content":" Let\u0026rsquo;s talk.\nSend a message Open for contract, consulting, and engineering work. Drop a note below or reach out directly at hanhphamit@gmail.com. Submitting opens a pre-addressed email draft in your local mail client: zero server endpoints, zero trackers.\nYour Name Subject Message Compose Email Draft prepared. If your email application did not open automatically, click here to send directly. Or connect directly: linkedin hanhphamphuoc email hanhphamit@gmail.com github hanhpp ","permalink":"https://hanhpham.vercel.app/contact/","summary":"\u003cdiv class=\"quantum-canvas-wrap\"\u003e\n  \u003ccanvas id=\"quantum-canvas\"\u003e\u003c/canvas\u003e\n\u003c/div\u003e\n\u003cp\u003eLet\u0026rsquo;s talk.\u003c/p\u003e\n\u003cdiv class=\"contact-form-wrap\"\u003e\n  \u003ch3\u003eSend a message\u003c/h3\u003e\n  \u003cp\u003eOpen for contract, consulting, and engineering work. Drop a note below or reach out directly at \u003ca href=\"mailto:hanhphamit@gmail.com\"\u003ehanhphamit@gmail.com\u003c/a\u003e. Submitting opens a pre-addressed email draft in your local mail client: zero server endpoints, zero trackers.\u003c/p\u003e\n  \u003cform class=\"contact-form\" id=\"contact-form\" onsubmit=\"return handleContactSubmit(event)\"\u003e\n    \u003cdiv class=\"contact-field\"\u003e\n      \u003clabel for=\"contact-name\"\u003eYour Name\u003c/label\u003e\n      \u003cinput type=\"text\" id=\"contact-name\" name=\"name\" placeholder=\"Optional\" maxlength=\"80\" autocomplete=\"name\"\u003e\n    \u003c/div\u003e\n    \u003cdiv class=\"contact-field\"\u003e\n      \u003clabel for=\"contact-subject\"\u003eSubject\u003c/label\u003e\n      \u003cinput type=\"text\" id=\"contact-subject\" name=\"subject\" placeholder=\"What is this about?\" required maxlength=\"120\"\u003e\n    \u003c/div\u003e\n    \u003cdiv class=\"contact-field\"\u003e\n      \u003clabel for=\"contact-message\"\u003eMessage\u003c/label\u003e\n      \u003ctextarea id=\"contact-message\" name=\"message\" placeholder=\"Write your message...\" required maxlength=\"2500\"\u003e\u003c/textarea\u003e\n    \u003c/div\u003e\n    \u003cbutton type=\"submit\" class=\"contact-submit-btn\"\u003e\n      \u003csvg xmlns=\"http://www.w3.org/2000/svg\" width=\"16\" height=\"16\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\u003e\u003cpath d=\"m22 2-7 20-4-9-9-4Z\"/\u003e\u003cpath d=\"M22 2 11 13\"/\u003e\u003c/svg\u003e\n      Compose Email\n    \u003c/button\u003e\n    \u003cdiv class=\"contact-status info\" id=\"contact-status\"\u003e\n      Draft prepared. If your email application did not open automatically, \u003ca id=\"contact-fallback-link\" href=\"#\"\u003eclick here to send directly\u003c/a\u003e.\n    \u003c/div\u003e\n  \u003c/form\u003e\n\u003c/div\u003e\n\u003cscript\u003e\nfunction handleContactSubmit(e) {\n    e.preventDefault();\n    var name = document.getElementById(\"contact-name\").value.trim();\n    var subject = document.getElementById(\"contact-subject\").value.trim();\n    var message = document.getElementById(\"contact-message\").value.trim();\n\n    if (!subject || !message) return false;\n\n    var body = message;\n    if (name) {\n        body = \"From: \" + name + \"\\n\\n\" + message;\n    }\n\n    var mailto = \"mailto:hanhphamit@gmail.com\"\n        + \"?subject=\" + encodeURIComponent(subject)\n        + \"\u0026body=\" + encodeURIComponent(body);\n\n    var statusBox = document.getElementById(\"contact-status\");\n    var fallbackLink = document.getElementById(\"contact-fallback-link\");\n    if (statusBox \u0026\u0026 fallbackLink) {\n        fallbackLink.href = mailto;\n        statusBox.classList.add(\"is-visible\");\n    }\n\n    window.location.href = mailto;\n    return false;\n}\n\u003c/script\u003e\n\u003ch3 id=\"or-connect-directly\"\u003eOr connect directly:\u003c/h3\u003e\n\u003cul class=\"contact-links\"\u003e\n  \u003cli\u003e\n    \u003cspan class=\"label\"\u003elinkedin\u003c/span\u003e\n    \u003ca href=\"https://www.linkedin.com/in/hanhphamphuoc/\" target=\"_blank\" rel=\"noopener\"\u003ehanhphamphuoc\u003c/a\u003e\n  \u003c/li\u003e\n  \u003cli\u003e\n    \u003cspan class=\"label\"\u003eemail\u003c/span\u003e\n    \u003ca href=\"mailto:hanhphamit@gmail.com\"\u003ehanhphamit@gmail.com\u003c/a\u003e\n  \u003c/li\u003e\n  \u003cli\u003e\n    \u003cspan class=\"label\"\u003egithub\u003c/span\u003e\n    \u003ca href=\"https://github.com/hanhpp\" target=\"_blank\" rel=\"noopener\"\u003ehanhpp\u003c/a\u003e\n  \u003c/li\u003e\n\u003c/ul\u003e\n\u003cscript\u003e\n(function () {\n    var canvas = document.getElementById(\"quantum-canvas\");\n    if (!canvas) return;\n    var ctx = canvas.getContext(\"2d\");\n    var wrap = canvas.parentElement;\n    var reduceMotion = window.matchMedia(\"(prefers-reduced-motion: reduce)\").matches;\n\n    function resize() {\n        var dpr = window.devicePixelRatio || 1;\n        canvas.width = wrap.clientWidth * dpr;\n        canvas.height = wrap.clientHeight * dpr;\n        ctx.scale(dpr, dpr);\n        canvas.style.width = wrap.clientWidth + \"px\";\n        canvas.style.height = wrap.clientHeight + \"px\";\n    }\n    resize();\n    window.addEventListener(\"resize\", resize);\n\n    var W = function () { return wrap.clientWidth; };\n    var H = function () { return wrap.clientHeight; };\n\n    var PARTICLE_COUNT = 18;\n    var CONNECT_DIST = 160;\n    var particles = [];\n\n    function getAccent() {\n        var s = getComputedStyle(document.documentElement);\n        return s.getPropertyValue(\"--accent\").trim() || \"#6d28d9\";\n    }\n\n    function hexToRgb(hex) {\n        hex = hex.replace(\"#\", \"\");\n        if (hex.length === 3) hex = hex[0]+hex[0]+hex[1]+hex[1]+hex[2]+hex[2];\n        return {\n            r: parseInt(hex.substring(0, 2), 16),\n            g: parseInt(hex.substring(2, 4), 16),\n            b: parseInt(hex.substring(4, 6), 16)\n        };\n    }\n\n    function Particle() {\n        this.x = Math.random() * W();\n        this.y = Math.random() * H();\n        this.vx = (Math.random() - 0.5) * 0.4;\n        this.vy = (Math.random() - 0.5) * 0.4;\n        this.r = 2 + Math.random() * 2;\n        this.phase = Math.random() * Math.PI * 2;\n        this.pulseSpeed = 0.01 + Math.random() * 0.02;\n    }\n\n    Particle.prototype.update = function () {\n        this.x += this.vx;\n        this.y += this.vy;\n        this.phase += this.pulseSpeed;\n        if (this.x \u003c 0) this.x = W();\n        if (this.x \u003e W()) this.x = 0;\n        if (this.y \u003c 0) this.y = H();\n        if (this.y \u003e H()) this.y = 0;\n    };\n\n    Particle.prototype.draw = function (accent) {\n        var pulse = 0.5 + 0.5 * Math.sin(this.phase);\n        var alpha = 0.4 + 0.4 * pulse;\n        var radius = this.r * (0.8 + 0.4 * pulse);\n\n        ctx.beginPath();\n        ctx.arc(this.x, this.y, radius, 0, Math.PI * 2);\n        ctx.fillStyle = \"rgba(\" + accent.r + \",\" + accent.g + \",\" + accent.b + \",\" + alpha + \")\";\n        ctx.fill();\n\n        ctx.beginPath();\n        ctx.arc(this.x, this.y, radius + 3, 0, Math.PI * 2);\n        ctx.strokeStyle = \"rgba(\" + accent.r + \",\" + accent.g + \",\" + accent.b + \",\" + (alpha * 0.3) + \")\";\n        ctx.lineWidth = 1;\n        ctx.stroke();\n    };\n\n    for (var i = 0; i \u003c PARTICLE_COUNT; i++) {\n        particles.push(new Particle());\n    }\n\n    function drawConnections(accent) {\n        for (var i = 0; i \u003c particles.length; i++) {\n            for (var j = i + 1; j \u003c particles.length; j++) {\n                var dx = particles[i].x - particles[j].x;\n                var dy = particles[i].y - particles[j].y;\n                var dist = Math.sqrt(dx * dx + dy * dy);\n                if (dist \u003c CONNECT_DIST) {\n                    var alpha = (1 - dist / CONNECT_DIST) * 0.25;\n                    ctx.beginPath();\n                    ctx.moveTo(particles[i].x, particles[i].y);\n                    ctx.lineTo(particles[j].x, particles[j].y);\n                    ctx.strokeStyle = \"rgba(\" + accent.r + \",\" + accent.g + \",\" + accent.b + \",\" + alpha + \")\";\n                    ctx.lineWidth = 1;\n                    ctx.stroke();\n                }\n            }\n        }\n    }\n\n    function animate() {\n        ctx.clearRect(0, 0, W(), H());\n        var accent = hexToRgb(getAccent());\n\n        drawConnections(accent);\n\n        for (var i = 0; i \u003c particles.length; i++) {\n            particles[i].update();\n            particles[i].draw(accent);\n        }\n\n        if (!reduceMotion) {\n            requestAnimationFrame(animate);\n        }\n    }\n\n    animate();\n})();\n\u003c/script\u003e","title":"Contact"},{"content":"In the previous post we self-hosted OSRM so distance calculations don\u0026rsquo;t depend on a metered maps API. OSRM answers in milliseconds even against a country-sized graph, but at real traffic volumes, the cheapest query is the one you don\u0026rsquo;t make at all. If your app repeatedly asks the distance between the same office and the same job site, or the same warehouse and the same few delivery zones, that\u0026rsquo;s a cache-hit rate waiting to be collected.\nThe instinct is to reach for something geospatial-native: PostGIS, a geohash library, an R-tree. For exact-match lookups on a bounded set of recurring routes, that\u0026rsquo;s more machinery than the problem needs. Plain Postgres, one rounding function, and a unique index cover it.\nThe 30-Second Pattern: Don\u0026rsquo;t install PostGIS just to cache recurring route lookups. Round incoming GPS coordinates to 4 decimal places (~11m precision), build a composite key (origin_lat, origin_lng, dest_lat, dest_lng), and back it with a standard Postgres UNIQUE B-tree index. You get 90%+ cache hit rates on repeat queries without geospatial extensions.\nThe key insight: round before you key GPS coordinates typically arrive with far more precision than the question needs. 13.723456789, 100.529099 and 13.723444444, 100.529050 are about a meter apart; for routing purposes, the same point. If you cache on raw float coordinates, you get a cache miss every time GPS jitter changes the 9th decimal place, which is close to always.\nRound to 4 decimal places instead: about 11 meters of precision, plenty for \u0026ldquo;which route is this\u0026rdquo;:\nfunc round4(v float64) float64 { return math.Round(v*10000) / 10000 } round4(13.723456789) = 13.7235 round4(13.723444444) = 13.7234 round4(100.529099) = 100.5291 round4(100.529050) = 100.5291 // ties round up, same as math.Round elsewhere That single function is doing the job a geohash would do (bucketing nearby points together) without adding a dependency or a new data type.\nThe schema CREATE TABLE IF NOT EXISTS distance_cache ( id BIGSERIAL PRIMARY KEY, from_lat NUMERIC(9,6) NOT NULL, from_lon NUMERIC(9,6) NOT NULL, to_lat NUMERIC(9,6) NOT NULL, to_lon NUMERIC(9,6) NOT NULL, distance_km NUMERIC(10,3) NOT NULL, created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(), UNIQUE (from_lat, from_lon, to_lat, to_lon) ); CREATE INDEX IF NOT EXISTS idx_distance_cache_lookup ON distance_cache (from_lat, from_lon, to_lat, to_lon); The UNIQUE constraint is doing double duty: it\u0026rsquo;s the cache key and it\u0026rsquo;s what makes an ON CONFLICT upsert possible, so a cache write is a single statement rather than a check-then-insert race.\nA composite unique index on four columns works here because a route is always queried as a specific ordered pair: from → to. If your app also needs to → from to hit the same cache row, round and sort the pair consistently before keying, or you\u0026rsquo;ll silently double your storage and your miss rate.\nReading and writing // Get returns (distanceKm, true) on cache hit, (0, false) on miss. func (s *Store) Get(ctx context.Context, fromLat, fromLon, toLat, toLon float64) (float64, bool, error) { fLat, fLon, tLat, tLon := round4(fromLat), round4(fromLon), round4(toLat), round4(toLon) var km float64 err := s.db.QueryRowContext(ctx, `SELECT distance_km FROM distance_cache WHERE from_lat=$1 AND from_lon=$2 AND to_lat=$3 AND to_lon=$4`, fLat, fLon, tLat, tLon, ).Scan(\u0026amp;km) if err == sql.ErrNoRows { return 0, false, nil } if err != nil { return 0, false, fmt.Errorf(\u0026#34;cache get: %w\u0026#34;, err) } return km, true, nil } // Set inserts or updates a cached distance. func (s *Store) Set(ctx context.Context, fromLat, fromLon, toLat, toLon, km float64) error { fLat, fLon, tLat, tLon := round4(fromLat), round4(fromLon), round4(toLat), round4(toLon) _, err := s.db.ExecContext(ctx, `INSERT INTO distance_cache (from_lat, from_lon, to_lat, to_lon, distance_km) VALUES ($1, $2, $3, $4, $5) ON CONFLICT (from_lat, from_lon, to_lat, to_lon) DO UPDATE SET distance_km=$5, created_at=NOW()`, fLat, fLon, tLat, tLon, km, ) return err } Wiring it in front of OSRM The handler tries the cache first and only falls through to OSRM on a miss, treating cache errors as non-fatal, since a cache that\u0026rsquo;s down should degrade to \u0026ldquo;slower,\u0026rdquo; never to \u0026ldquo;broken\u0026rdquo;:\nfunc (h *Handler) Distance(w http.ResponseWriter, r *http.Request) { var req distanceRequest if err := json.NewDecoder(r.Body).Decode(\u0026amp;req); err != nil { writeErr(w, http.StatusBadRequest, \u0026#34;invalid JSON\u0026#34;) return } ctx := r.Context() km, hit, err := h.cache.Get(ctx, req.From.Lat, req.From.Lon, req.To.Lat, req.To.Lon) if err != nil { log.Printf(\u0026#34;cache get error: %v\u0026#34;, err) // non-fatal: fall through to OSRM } if hit { writeJSON(w, distanceResponse{DistanceKm: km, Cached: true}) return } km, err = h.osrm.RoadDistanceKm(ctx, req.From.Lat, req.From.Lon, req.To.Lat, req.To.Lon) if err != nil { writeErr(w, http.StatusBadGateway, \u0026#34;could not calculate distance\u0026#34;) return } if err := h.cache.Set(ctx, req.From.Lat, req.From.Lon, req.To.Lat, req.To.Lon, km); err != nil { log.Printf(\u0026#34;cache set error: %v\u0026#34;, err) // non-fatal } writeJSON(w, distanceResponse{DistanceKm: km, Cached: false}) } That Cached field in the response isn\u0026rsquo;t just debug noise: exposing it makes cache-hit rate observable from the outside without needing to add metrics plumbing before you\u0026rsquo;ve decided whether you need any.\nWhere this approach stops working Rounded-coordinate caching is a good fit specifically because the traffic is exact-match repeats: the same handful of real-world places, queried over and over. It stops being the right tool once the questions change shape:\n\u0026ldquo;Points within N km of here\u0026rdquo;: a proximity/radius query needs an actual spatial index (PostGIS GIST on a geography column, or equivalent); a unique index on rounded columns can\u0026rsquo;t answer \u0026ldquo;nearby,\u0026rdquo; only \u0026ldquo;identical.\u0026rdquo; Unbounded, low-repeat coordinates: if every query is between two points that have never been queried before (e.g. live GPS breadcrumbs), cache hit rate approaches zero and the table just grows forever with no benefit. Add a TTL/eviction policy, or don\u0026rsquo;t cache at all. Sub-11-meter precision requirements: round4 buys ~11m buckets; tighten the rounding if your use case needs finer-grained distinction between very close points, at the cost of more distinct cache rows. For the common case this was built for (recurring routes between a bounded set of known locations), a table, a rounding function, and a unique index outperform the complexity of a geospatial engine you\u0026rsquo;d otherwise have to run and operate for the same result.\n","permalink":"https://hanhpham.vercel.app/posts/caching-geospatial-queries-without-a-geo-database/","summary":"You don\u0026rsquo;t need PostGIS or a geohash library to cache repeat road-distance lookups: rounding coordinates and a unique index gets you most of the win.","title":"Caching Geospatial Queries Without a Specialized Geo Database"},{"content":"If your product calculates distance between two coordinates (a delivery fee, a travel expense claim, a \u0026ldquo;drivers near you\u0026rdquo; radius), the tempting default is a maps API: send two points, get back a number, pay per request. That works, but it puts a metered, rate-limited third party in the critical path of a calculation you might run millions of times.\nThe alternative is a fully open-source routing stack you run yourself: OpenStreetMap data, preprocessed by OSRM, served from a container you control. No API key, no per-request billing, no external outage taking down your pricing engine.\nWhy not just use Haversine? Straight-line (Haversine) distance is free and requires no infrastructure at all; it\u0026rsquo;s just trigonometry on two lat/lon pairs. The problem is that roads aren\u0026rsquo;t straight lines. Rivers, one-way systems, and the fact that Bangkok is not a grid all mean actual travel distance is consistently longer than the straight-line distance between two points:\nFROM TO ROAD KM HAVERSINE KM RATIO ──── ── ─────── ──────────── ───── Silom BTS Siam BTS 3.33 2.59 ×1.29 Asok Victory Monument 5.11 4.00 ×1.28 Don Mueang Airport Suvarnabhumi Airport 39.30 29.18 ×1.35 Nonthaburi (suburban) Silom BTS 20.97 15.52 ×1.35 Bangkok routes average around ×1.3 (road distance a third longer than straight-line), and that ratio isn\u0026rsquo;t constant, it depends on the specific geography between two points. There\u0026rsquo;s no fixed multiplier you can apply to Haversine and get something trustworthy; you have to actually route.\nIf a distance figure feeds into money (reimbursement, delivery pricing, SLA radius), Haversine will be wrong in a way that\u0026rsquo;s hard to justify later. It\u0026rsquo;s fine for \u0026ldquo;how far away, roughly\u0026rdquo; UI, not fine for a number on an invoice.\nThe stack Component Role OpenStreetMap (OSM) Free, community-maintained map data: roads, turn restrictions, speed limits Geofabrik Hosts daily per-country OSM extracts as .osm.pbf files OSRM Preprocesses OSM data into a routing graph, answers distance/duration queries in milliseconds You need Docker and a .osm.pbf extract for whatever region you\u0026rsquo;re routing in. Nothing here calls out to a paid service.\nGetting the map data Geofabrik publishes extracts per country. For Thailand:\nmkdir -p osm-data curl -L -o osm-data/thailand-latest.osm.pbf \\ https://download.geofabrik.de/asia/thailand-latest.osm.pbf This file is roughly 300–400 MB: most of a country\u0026rsquo;s entire road network, free to download, no account required.\nPreprocessing: extract, partition, customize OSRM doesn\u0026rsquo;t route directly against the raw PBF; it has to build a routing graph first, in three steps. This is the part that actually costs CPU time (a few minutes), and it only needs to happen once per map version:\nservices: osrm-init: image: ghcr.io/project-osrm/osrm-backend:latest volumes: - ./osm-data:/pbf:ro - osrm_data:/data entrypoint: /bin/sh command: - -c - | set -e osrm-extract -p /opt/car.lua /pbf/thailand-latest.osm.pbf -o /data/thailand.osrm osrm-partition /data/thailand.osrm osrm-customize /data/thailand.osrm osrm: image: ghcr.io/project-osrm/osrm-backend:latest volumes: - osrm_data:/data command: osrm-routed --algorithm mld /data/thailand.osrm ports: - \u0026#34;5000:5000\u0026#34; restart: unless-stopped depends_on: osrm-init: condition: service_completed_successfully volumes: osrm_data: osrm-extract reads the PBF and a routing profile (car.lua ships with the image) to build the base graph; this is where \u0026ldquo;which roads can cars use, and at what speed\u0026rdquo; gets decided. osrm-partition and osrm-customize prepare the graph for the mld (multi-level Dijkstra) query algorithm, which is what makes queries against a country-sized graph return in milliseconds instead of seconds. The osrm service only starts once osrm-init exits successfully, so a single docker compose up -d handles first-run preprocessing and every subsequent restart correctly; restarts skip straight to serving, since the processed graph is sitting in the osrm_data volume. docker compose up -d docker compose logs -f osrm-init # watch preprocessing progress OSRM is ready when it\u0026rsquo;s listening on http://localhost:5000.\nQuerying it OSRM\u0026rsquo;s HTTP API takes coordinates in lon,lat order: the GeoJSON convention, and the opposite of the lat,lon order most map UIs use. This is the single easiest mistake to make integrating against it:\ncurl \u0026#34;http://localhost:5000/route/v1/driving/100.5352,13.7563;100.5018,13.7466?overview=false\u0026#34; A minimal Go client wrapping that endpoint:\npackage osrm import ( \u0026#34;context\u0026#34; \u0026#34;encoding/json\u0026#34; \u0026#34;fmt\u0026#34; \u0026#34;net/http\u0026#34; \u0026#34;time\u0026#34; ) type Client struct { base string http *http.Client } func NewClient(baseURL string) *Client { return \u0026amp;Client{base: baseURL, http: \u0026amp;http.Client{Timeout: 10 * time.Second}} } type routeResponse struct { Code string `json:\u0026#34;code\u0026#34;` Routes []struct { Distance float64 `json:\u0026#34;distance\u0026#34;` // meters Duration float64 `json:\u0026#34;duration\u0026#34;` // seconds } `json:\u0026#34;routes\u0026#34;` } // RoadDistanceKm returns the shortest road distance in kilometers between two points. func (c *Client) RoadDistanceKm(ctx context.Context, fromLat, fromLon, toLat, toLon float64) (float64, error) { url := fmt.Sprintf( \u0026#34;%s/route/v1/driving/%f,%f;%f,%f?overview=false\u0026amp;alternatives=false\u0026#34;, c.base, fromLon, fromLat, toLon, toLat, ) req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil) if err != nil { return 0, err } resp, err := c.http.Do(req) if err != nil { return 0, fmt.Errorf(\u0026#34;osrm request: %w\u0026#34;, err) } defer resp.Body.Close() if resp.StatusCode != http.StatusOK { return 0, fmt.Errorf(\u0026#34;osrm returned status %d\u0026#34;, resp.StatusCode) } var result routeResponse if err := json.NewDecoder(resp.Body).Decode(\u0026amp;result); err != nil { return 0, fmt.Errorf(\u0026#34;osrm decode: %w\u0026#34;, err) } if result.Code != \u0026#34;Ok\u0026#34; || len(result.Routes) == 0 { return 0, fmt.Errorf(\u0026#34;osrm: no route found (code=%s)\u0026#34;, result.Code) } return result.Routes[0].Distance / 1000.0, nil } overview=false skips returning the full route geometry, worth setting explicitly if all you need is the distance number, since the polyline for a 40km route is a lot of payload you\u0026rsquo;d otherwise discard.\nWhat this costs instead of money Nothing per-request, but it isn\u0026rsquo;t free in the way \u0026ldquo;no external dependency\u0026rdquo; sounds:\nDisk and memory for the processed graph: a whole-country graph is sized in gigabytes, not megabytes. Map freshness is on you: Geofabrik extracts update daily, but your running graph is frozen until you re-run extract/partition/customize against a newer PBF. You own the failure mode: if the container falls over, that\u0026rsquo;s your on-call, not a status page you can just check. For a service making a meaningful volume of distance queries, that trade is usually worth it. For an occasional lookup, a paid API\u0026rsquo;s simplicity may still win; the point isn\u0026rsquo;t that self-hosting is always right, it\u0026rsquo;s that it\u0026rsquo;s a real option once volume or cost makes the metered version uncomfortable.\nNext: don\u0026rsquo;t ask OSRM the same question twice Preprocessing gets queries down to milliseconds, but at real traffic volumes (the same handful of office-to-site routes queried over and over for expense claims, say), even a millisecond round-trip adds up, and there\u0026rsquo;s no reason to recompute a route that hasn\u0026rsquo;t changed. The next post covers caching these queries in a plain Postgres table, without reaching for a specialized geo database.\n","permalink":"https://hanhpham.vercel.app/posts/self-hosting-a-routing-engine/","summary":"Road distance, not straight-line distance, is what travel expense claims and delivery pricing actually need, here\u0026rsquo;s how to get it without a metered maps API.","title":"Self-Hosting a Routing Engine Instead of Paying Per API Call"},{"content":"This post lives in its own folder (content/posts/using-images-and-shortcodes/index.md), alongside an image file. That folder is a Hugo page bundle: any asset placed next to index.md can be referenced with a relative path, and it gets copied to the right place automatically when the site builds.\nThe figure shortcode {{\u0026lt; figure src=\u0026#34;cover.svg\u0026#34; alt=\u0026#34;Demo cover image\u0026#34; caption=\u0026#34;A generated SVG, bundled with this post\u0026#34; \u0026gt;}} renders as:\nA generated SVG, bundled with this post\nWhy bundles are handy Move the folder, and the post keeps its images; no broken links. No separate static/images/... path to keep in sync with the post. Works for any file type: PDFs, diagrams, downloadable code samples. Linking to other posts Shortcodes like {{\u0026lt; ref \u0026gt;}} resolve internal links by filename instead of a hardcoded URL, so renaming a post\u0026rsquo;s slug doesn\u0026rsquo;t quietly break links from other pages; see the welcome post for an example.\n","permalink":"https://hanhpham.vercel.app/posts/using-images-and-shortcodes/","summary":"Keeping a post and its images together with Hugo page bundles, plus the built-in figure shortcode.","title":"Page Bundles, Images, and Shortcodes"},{"content":"One thing I like about Hugo is that fenced code blocks get syntax highlighting for free, powered by Chroma. Here are a few languages in action.\nGo package main import \u0026#34;fmt\u0026#34; func fib(n int) int { if n \u0026lt; 2 { return n } a, b := 0, 1 for i := 2; i \u0026lt;= n; i++ { a, b = b, a+b } return b } func main() { fmt.Println(fib(10)) // 55 } Python def fib(n: int) -\u0026gt; int: a, b = 0, 1 for _ in range(n): a, b = b, a + b return a print([fib(i) for i in range(10)]) Shell # Build the site and start a local preview server hugo server --buildDrafts --disableFastRender A quick table too Language Paradigm First appeared Go Compiled 2009 Python Interpreted 1991 Bash Shell scripting 1989 And inline code like hugo new content posts/my-post.md works anywhere in a sentence.\n","permalink":"https://hanhpham.vercel.app/posts/hugo-code-blocks-demo/","summary":"Hugo ships with built-in syntax highlighting via Chroma, no plugins required.","title":"Syntax Highlighting with Hugo"},{"content":"Hi, I\u0026rsquo;m Hanh. This is the first post on a small blog I put together with Hugo, a static site generator written in Go, and the PaperMod theme.\nWhy a static blog? A few reasons I like this setup over a hosted platform:\nFast: pages are pre-rendered HTML, no server-side rendering per request. Free hosting: GitHub Pages serves it directly from this repo. Markdown-native: posts are just .md files under content/posts/, so writing and version control both stay simple. No lock-in: the content lives in plain text, not a proprietary database. What to expect This is a demo site, so the next couple of posts show off some of Hugo\u0026rsquo;s content features:\nCode blocks and syntax highlighting Images and shortcodes More real posts to come. Thanks for stopping by!\n","permalink":"https://hanhpham.vercel.app/posts/welcome-to-my-blog/","summary":"A quick hello and a look at how this site is put together.","title":"Welcome to My Blog"}]