{
  "archive": {
    "site": "https://vihaanvaghela.com",
    "generated": "2026-08-07T18:59:59.143Z",
    "format": 1,
    "note": "The complete public archive. Drafts and anything marked private are absent. Bodies are Markdown. Every entry carries both the date it happened and the date it was recorded, its corrections, and whether a machine helped write it. Copy it, mirror it, keep it — that is what it is for.",
    "counts": {
      "essays": 1,
      "documents": 34,
      "projects": 4,
      "milestones": 19,
      "corrections": 2
    },
    "checksum": "cd7510320ed55bf1d3ac9d272dcba7ca83acedd3252be59023e0c7d5796c5be9"
  },
  "entries": [
    {
      "kind": "essay",
      "slug": "print-hello-world",
      "title": "print(\"Hello, World!\")",
      "summary": "The first post. The start to growth.",
      "body": "Hello. By now, you probably know who I am. My name is Vihaan Vaghela. After all, this website exists to document my thoughts and the systems I hope will outlast their creator. This is my first post. Now, I haven't fully explained why I decided to post in the first place. Let me explain:\n\nSee my parents always forced me to write a journal. I was inconsistent. Then I realized the so called \"importance\" of college. I started thinking about college. I thought a lot. I researched a ton, and I later found out that to get into college, grades (which I thought were the only factor) are not merely enough. Admission officers look at extracurriculars, achievements, and various other things. Leadership and an impact on the community by expressing your views, was a key part. \n\nThis is how I found out, that when my parents told me to journal, blog, vlog and express my thoughts, they were looking out for me and not forcing me to just mindlessly write. \n\nThen I started to journal. I was consistent, until I wasn't...\nConsistency was the key to the lock to the doors of success. But whenever I tried to 'pick' that lock, time was my biggest enemy. I got distracted. I used to think that one day without journaling won't hurt. Will it?\nLong story short, it did. \nI didn't journal for 7 months on end and one day when I was lying in bed, scrolling, I became victim of a canonical event that almost every man is bound to go through. That event is now known as \"THE GOGGINS EFFECT\". Derived from the legendary David Goggins. It made me realize that I had potential, but I was wasting all of it by coming home, eating, playing Brawlstars. I hated myself. I decided to change. I did change. I changed a lot. \n\n\nThis is what compelled me to start blogging and expressing my views. Not college (though it was a major revelation inducing factor), not Goggins, but evolution. Growth.\n\nMy blogs are not going to be perfect. I'm sure that someplace or the other, you will be able to find an error, but that's not the point. It's about growth. All of my writings on this tab will be purely my own words. Not a single excerpt of an AI written text will be mentioned on this tab. I will document anything and everything that I feel like deserves a post, and in some cases I might also dedicate a entire post to a rant about something bothering me 😂.\n\nI started blogging because I wanted to grow. And grow I will. \nWelcome to the archives.\n\n-Vihaan Vaghela.",
      "status": "PUBLISHED",
      "occurred": "2026-08-03T17:28:37.838Z",
      "recorded": "2026-08-03T17:28:37.838Z",
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/blog/print-hello-world"
    },
    {
      "kind": "document",
      "slug": "adaptive-systems-over-static-systems",
      "title": "Adaptive systems over static systems",
      "summary": "Why the preference for adaptation is conditional rather than absolute, and what a system must have earned before it is allowed to adapt.",
      "body": "This principle is the one most easily misread, because stated baldly it sounds like a\npreference for sophistication. It is not. It is a claim about which failure is\npreferable, and it comes with conditions that most adaptive systems do not meet.\n\n**The principle: adaptation is preferred to a fixed schedule — but only for a system\nthat can detect when its adaptation is no longer justified, and stop.**\n\n## The argument for adaptation\n\nA static system encodes a decision made once and applies it indefinitely. Its central\ndefect is not that the decision was poor; it is that the system has no mechanism for\ndiscovering that conditions have moved. It cannot be wrong in a way it can notice,\nwhich means it cannot improve and cannot alert.\n\nAn adaptive system can be wrong in ways it can notice. That is the whole of the\nadvantage, and it is a large one — but notice that the advantage is about detection,\nnot about accuracy. A system that adapts and cannot evaluate its own adaptation has\ntaken on the additional failure modes and none of the benefit.\n\n## The conditions\n\nAdaptation is justified only when the system can do three things, and a system missing\nany of them is better off static.\n\n**It must know when it is outside its competence.** Conditions leave the range a model\nwas built for. If that departure is not detectable, the system produces confident\noutput in exactly the circumstances where confidence is unwarranted.\n\n**It must have somewhere to go.** Detecting incompetence is only useful if there is a\ndefined behaviour to fall back to — see\n[governance before intelligence](/vector/governance-before-intelligence).\n\n**It must be able to account for what it did.** An adaptive system that cannot explain\na decision cannot be corrected after a bad one, so its adaptations accumulate without\never being audited.\n\n## Why the preference is not absolute\n\nThere are components in VECTOR that are deliberately static, and they are the ones\nthat matter most in a crisis. Emergency shutdown does not adapt. Termination does not\nadapt. Their behaviour is fixed, simple, and reasoned about completely.\n\nThe rule that resolves this: **adapt where being wrong is recoverable; be static where\nit is not.** Sophistication is appropriate in proportion to how reversible the\nconsequences are. An adaptive emergency stop is a contradiction — the entire value of\nthe mechanism is that its behaviour is known in advance, including under conditions\nnobody anticipated.\n\n## Failure mode\n\nThe characteristic failure is adaptation during instability. A system that adjusts its\nown parameters while conditions are deteriorating amplifies exactly what it is trying\nto damp, and the resulting oscillation is difficult to diagnose because every\nindividual adjustment was locally reasonable.\n\nThe defence is that adaptation is itself gated: the slowest components in the system\nare the ones that change it, and they decline to act while conditions are stressed.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/adaptive-systems-over-static-systems"
    },
    {
      "kind": "document",
      "slug": "adversarial-robustness",
      "title": "Adversarial robustness",
      "summary": "Statistical anomaly detection assumes sensors are noisy. This subsystem exists for the case where they are lying.",
      "body": "*[VECTOR archive](/vector) › Architecture › Adversarial robustness*\n\n[Sentinel](/vector/protocol-sentinel) detects values that are statistically unusual.\nThat is the correct defence against a sensor that is broken. It is no defence at all\nagainst a sensor that is being manipulated, because manipulated inputs are constructed\nto be unremarkable.\n\n## Problem\n\nInfrastructure that responds to observation can be influenced by anyone who can\ninfluence the observation. An adversary does not need access to the decision system if\nthey can feed it plausible, false readings — and plausibility is exactly what defeats\nthreshold and z-score detection.\n\n## Context\n\nThis was identified as a gap before it was addressed, and named as future work in the\npublished account of v1. It now exists as a subsystem, which is worth stating plainly:\nthe roadmap item has been built.\n\n## Constraints\n\nIt has to operate on the same inputs the live system consumes, cheaply enough to sit in\na real-time path, and without a labelled corpus of attacks — because the attacks that\nmatter are the ones nobody has catalogued.\n\n## Alternatives considered\n\n**Wider statistical thresholds.** Catches cruder manipulation and, by construction,\nmisses anything shaped to look normal. It also raises false positives, which erodes\nconfidence in the whole chain.\n\n**Supervised attack classification.** Effective against known attack classes, and it\nrequires examples of attacks. A defence that only recognises what it has been shown is\nnot a defence against a deliberate adversary.\n\n## Chosen architecture\n\n`security/adversarial_robustness.py`, built around a `RobustnessAutoencoder` with a\n`RobustnessTrainer`, `RobustnessConfig` and `RobustnessTrainingConfig`, exposed at\nruntime as an `AdversarialRobustnessGuard`.\n\nThe autoencoder approach answers the labelled-data constraint. An autoencoder is trained\nto reconstruct normal inputs and is scored by how badly it fails: an input drawn from\nthe distribution it learned reconstructs well, and an input constructed adversarially\nreconstructs poorly, because it is not from that distribution however plausible its\nindividual values look.\n\nThis detects unfamiliarity rather than abnormality, which is a genuinely different\nquestion from the one Sentinel asks — and the reason both exist.\n\n`SecurityDataset` and `collate_security_samples` provide the training path; the `Guard`\nis the runtime boundary.\n\n## Subsystem relationships\n\nSits alongside detection rather than inside it. Where Sentinel asks \"is this value\nunusual for this system\", the guard asks \"does this input resemble anything I was\ntrained on\". Its output belongs upstream of the confidence that\n[observation before action](/vector/observation-before-action) requires every\nobservation to carry.\n\n## Data flow\n\n```\ninput → RobustnessAutoencoder → reconstruction error → AdversarialRobustnessGuard\n                                                     → trust signal on the observation\n```\n\n## Tradeoffs\n\n**Legitimate novelty is indistinguishable from attack.** A genuinely unprecedented but\nreal condition reconstructs poorly, for the same reason a crafted input does. The guard\ncannot tell them apart, and both narrow autonomy — the safe direction, and a real cost\nduring unusual legitimate events.\n\n**The defence inherits its training distribution.** An autoencoder trained on a period\nof normality treats drift away from that period as suspicious.\n\n**A learned defence has its own adversarial surface.** An adversary who knows the\nmechanism can, in principle, construct inputs that reconstruct well and are still\nfalse. This is an honest limitation of the approach, not a defect of this implementation.\n\n## Failure modes\n\n**False positives during real anomalies** — a legitimate emergency is unusual, and\nunusual is what the guard reports on.\n\n**Silent staleness.** A guard whose training distribution no longer describes the\nsystem degrades gradually and without announcing it.\n\n## Future evolution\n\nPeriodic retraining against recent normal operation is the obvious requirement, and it\ntrades directly against the guard being able to notice slow drift — the same tension\nSentinel has with its rolling window. Whether the current implementation retrains is not\nestablished by structure alone.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/adversarial-robustness"
    },
    {
      "kind": "document",
      "slug": "changelog-v1",
      "title": "v1 — 647,000 cycles, zero failures",
      "summary": "The first working v1: a two-hour adversarial stress test with no failures, no crashes, no dropped events and no ordering violations.",
      "body": "v1 is real.\n\n## The test\n\nA two-hour adversarial stress test across more than 647,000 total cycles:\n\n- Zero failures\n- Zero crashes\n- Zero dropped events on the bus\n- Zero ordering violations\n- 6.17 ms p95 latency under sustained adversarial load\n\nThe hardware abstraction layer is production-ready. The backend API is live. The\ndecision engine is proven.\n\n## What shipped\n\nThe full eight-protocol governance stack, from Sentinel through to Terminus. The\nintelligence layer — Anarchy, ORACLE, MORL, PULSAR — coordinated by Fusion under\nsafety gates, with the Meta controller adapting slowly above it. The authority\nprotocol with its deterministic escalation and staged recovery.\n\nAnd the thing that made all of it possible: per-stage time budgets, enforced, after\n[the latency crisis](/vector/the-latency-crisis).\n\n## What is not done\n\nSUMO training and a real pilot. Until a system has met an environment it did not\nanticipate, its numbers describe a laboratory.\n\n## Roadmap\n\n- **Temporal graph networks** — extend graph modelling to how the graph itself evolves\n  over time, rather than treating each snapshot independently\n- **Digital twin engine** — a real-time virtual mirror of the network running alongside\n  the live system rather than ahead of it\n\n## Since shipped\n\nThree items published on the v1 roadmap have since been built, and this entry is\namended rather than rewritten so the sequence stays visible:\n\n- **Graph neural network traffic modelling** now exists as `learning/gnn_traffic_model.py`\n  and `learning/gnn_prediction_layer.py` — message passing over the road network as a\n  graph, with its own training and inference paths.\n- **Adversarial robustness** exists as `security/adversarial_robustness.py`. It is\n  documented in [adversarial robustness](/vector/adversarial-robustness).\n- **Predictive incident detection** exists as `learning/incident_detection.py`.\n\nA changelog whose roadmap silently becomes accurate is a changelog nobody can date.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/changelog-v1"
    },
    {
      "kind": "document",
      "slug": "documentation-is-part-of-engineering",
      "title": "Documentation is part of engineering",
      "summary": "Writing the argument for a component before building it is the cheapest way to discover that it should not exist.",
      "body": "Documentation is conventionally understood as a description of a system, produced\nafter the system exists, for the benefit of people who did not build it. Under that\ndefinition it is hygiene — worthwhile, secondary, and the first thing dropped under\nschedule pressure.\n\n**The principle: a component requires a written argument for its existence before it\nis built, and the argument is engineering work rather than a description of it.**\n\n## What the argument has to contain\n\nNot a description of what the component will do. An argument for why it should exist:\nwhat it is responsible for, why that responsibility cannot live in something that\nalready exists, what alternatives were considered and why they lost, and what would\nhave to be true for this to be the wrong idea.\n\nThe last item is the one that does the work. A justification that cannot state its own\nfalsification condition is usually a preference that has been written in the register\nof a reason.\n\n## Why it belongs before the code\n\nBecause a meaningful fraction of the time, writing it is how you discover the component\nshould not exist. The responsibility turns out to belong to something already present,\nor the problem it addresses turns out to be a symptom of a problem one layer down.\n\nThe economics are lopsided enough that the practice pays for itself even at a low hit\nrate: an hour of writing against a week of building, the ongoing maintenance of\nsomething that should never have existed, and the considerably larger cost of only\ndiscovering the mistake once the architecture has grown around it.\n\nDocumentation written afterwards cannot do this. By then the decisions are made and the\nhonest options are to describe what exists or to misrepresent it. So it describes, and\nthe description is accurate and nearly worthless, because the part a future reader\nneeds is the part that is gone — the alternatives, and why they lost.\n\n## The symmetry worth noticing\n\nThis is the same argument as\n[explainability before automation](/vector/explainability-before-automation), applied\nto people instead of components. An explanation generated after the fact by a separate\nprocess is a reconstruction, whether the process is a neural network or an engineer\nwriting a design document from memory. In both cases the account must be produced by\nthe same activity that produced the decision, or it is a story about the decision.\n\n## The failure mode\n\nDocuments are pleasant to write. They feel like progress and cost nothing to be wrong\nabout, and it is entirely possible to spend a week architecting something that two\ndays of building would have settled empirically.\n\nThe boundary in use: **write the document when the cost of being wrong is structural.**\nAnything other components will depend on, anything expensive to reverse, anything\ntouching the authority boundary — document first. Anything that can be deleted in an\nafternoon — build it and find out.\n\nErring toward writing is slow. Erring the other way produces a system nobody decided\non, which is not recoverable by working harder.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/documentation-is-part-of-engineering"
    },
    {
      "kind": "document",
      "slug": "explainability-before-automation",
      "title": "Explainability before automation",
      "summary": "A component that cannot account for its output does not enter the decision path, regardless of how well it performs.",
      "body": "There is a standard order of operations in automated systems: build something that\nworks, then add an explanation layer so that people can follow it. The order is\nbackwards, and the reason is not ethical.\n\n**The principle: a component that cannot account for its output does not enter the\ndecision path, however well it performs.**\n\n## Why the usual order fails\n\nAn explanation produced after a decision, by a different mechanism than the one that\nproduced the decision, is a reconstruction. It has access to the inputs and the output\nand it infers a plausible path between them.\n\nSometimes the plausible path is not the actual path. And the failure is undetectable\nfrom outside: a reconstructed explanation and a truthful one look identical, and the\nonly way to catch a discrepancy is to already know what the system did — which is what\nthe explanation was for.\n\nThe result is a component that is fluent, confident, and occasionally lying about the\nsystem it describes. That is worse than no explanation, because no explanation at\nleast announces the absence.\n\n## What the principle requires instead\n\nThe justification must be produced by the same process that produced the decision, in\nthe same pass, from the same values. It is part of the decision's structure rather\nthan a report about it — the difference between a witness and a narrator.\n\nA decision therefore carries the observations it rested on and their confidences, the\nprediction it trusted and how far, the constraints that were active at that moment,\nand the counterfactual: what it would have done had any of those differed.\n\nThe counterfactual is what makes the record useful rather than merely honest. \"The\nsignal held because queue length exceeded threshold\" is a description. Adding \"and\nwould have released at a queue length two lower, or had the emergency constraint been\nactive\" produces something a person can argue with.\n\n## The cost, stated plainly\n\nThis principle disqualifies methods. A component whose output cannot be accounted for\nmay not sit in the decision path, no matter how much better it scores, and this is not\na preference to be traded away under deadline pressure.\n\nThe justification is the same one that runs through\n[why VECTOR exists](/vector/why-vector-exists): a system that performs well and cannot\naccount for itself cannot be corrected, cannot be audited, and cannot be defended when\nit was right. Its good performance is not knowledge. It is luck that has not yet run\nout, and nobody can tell the difference from the outside.\n\n## Relationship to the previous principle\n\nExplainability is only possible because of\n[observation before action](/vector/observation-before-action). A decision can record\nwhat it rested on only if what it rested on arrived with its provenance intact. The\ntwo principles are one requirement observed at two points in the pipeline.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/explainability-before-automation"
    },
    {
      "kind": "document",
      "slug": "governance-before-intelligence",
      "title": "Governance before intelligence",
      "summary": "The structure that constrains a capability must exist before the capability does. Retrofitting authority onto a working system does not work.",
      "body": "The natural order of construction is to build the thing that does the work, confirm it\nworks, and then add the structure that keeps it safe. This order is close to universal\nand it is wrong, for a reason that is structural rather than a matter of discipline.\n\n**The principle: the structure that governs a capability is built and proven before the\ncapability exists.**\n\n## Why retrofitting fails\n\nGovernance added to a working system can only wrap it. The system already produces its\noutputs through paths that were designed without reference to constraint, so the\nconstraint has to sit outside and reject what it does not like.\n\nThat produces a veto, and a veto has a specific defect: the reasoning that produced the\nrejected action is wasted, and the system now needs a fallback for a situation it did\nnot anticipate. Worse, vetoes fire most often when conditions are unusual — which is\nprecisely when the unreasoned fallback is least appropriate.\n\nA constraint present *during* the decision behaves differently. It is a boundary the\nsystem reasons inside, so a disallowed action is never produced, and there is no\nfallback path because there is nothing to fall back from. This distinction cannot be\nretrofitted; it is a property of where the constraint lives.\n\n## The second reason: the floor\n\nBuilding governance first also means building a system that works without any\nintelligence at all — a deterministic controller whose behaviour can be reasoned about\ncompletely.\n\nThat controller is not scaffolding to be discarded. It is the floor. It gives every\nlearned component something to beat, which converts \"the model is better than nothing\"\ninto a measurable claim. And it gives every degradation path somewhere defined to land,\nso that when intelligence is gated off the system does not enter an undefined state —\nit returns to the thing that was there first.\n\nA system without a floor faces a choice, when it is uncertain, between doing something\nclever and doing nothing. Both are bad, and the choice is unnecessary.\n\n## What this costs\n\nIt is a genuinely unsatisfying way to work. Months spent on containment, escalation and\ntermination produce no benchmark improvement whatsoever. There is nothing to show. The\nsystem does not appear more capable at the end of it than at the beginning.\n\nIt is also the reason the project survived its own scope change. When the domain\nexpanded, the governance structure did not have to move, because it had never been\nshaped around the intelligence it was constraining.\n\n## Corollary\n\nCapability is admitted incrementally and can be withdrawn. A component enters the\ndecision path only after the structure that constrains it exists, and remains subject\nto being gated off — under low confidence, high risk, or authority intervention —\nwithout the system losing definition. See\n[the human remains the operator](/vector/the-human-remains-the-operator).",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/governance-before-intelligence"
    },
    {
      "kind": "document",
      "slug": "observation-before-action",
      "title": "Observation before action",
      "summary": "The first principle. A system may not act on a quantity it has not observed, and an observation is incomplete without its uncertainty.",
      "body": "Every control system contains an implicit claim about the state of the world. The\nquestion is whether that claim was measured or assumed, and whether the system can\ntell the difference.\n\n**The principle: a component may not act on a quantity it has not observed, and an\nobservation is not complete until it carries how much it should be trusted.**\n\n## The first half\n\nThe prohibition is on acting from assumption where observation was available. A fixed\nschedule is the pure case — every action rests on a state nobody checked — but the\nmore common violation is subtler: a component that observes one quantity and infers\nanother without recording that it inferred it.\n\nThe inference may be sound. The problem is that it becomes indistinguishable from a\nmeasurement one layer downstream, and any later component reasoning about reliability\nnow has bad information about its own inputs.\n\n## The second half, which is the one that matters\n\nA measurement without its uncertainty is a number pretending to be a fact.\n\nTwo sensors reporting the same value are not equivalent if one is reporting from a\nwell-understood regime and the other is extrapolating past its calibrated range. A\nsystem given only the values cannot distinguish them, and will treat both as equally\nactionable — which means it behaves identically whether it knows what is happening or\nnot.\n\nCarrying confidence alongside every observation is what makes every downstream safety\nproperty expressible. A component can decline to act under low confidence only if\nconfidence reached it. An arbitration layer can weight sources only if the weights\nhave a basis. An authority layer can narrow autonomy under uncertainty only if\nuncertainty is a value it can read.\n\nStrip the uncertainty at the boundary and none of those are available later, at any\nprice.\n\nThe principle also constrains what may be observed at all. A system that admits every\navailable signal has admitted a great many that stop being informative under exactly\nthe conditions where they matter — see\n[how a complex system was reduced to seven variables](/vector/seven-variables).\n\n## The cost\n\nIt is not free. Every interface widens, every component must decide what to do with a\nconfidence it might rather ignore, and there is a persistent temptation to collapse\nthe pair back into a single number \"for now\".\n\nThat collapse is irreversible in practice. The information is gone at the point it is\ndiscarded, and reconstructing it afterwards means inventing it.\n\n## Failure mode\n\nThe characteristic failure is confident action on degraded input, and it is silent.\nThe system does not behave erratically — it behaves normally, on data that no longer\nmeans what it did. Nothing alerts, because from the inside the values are within range.\n\nThis is why the principle is first. It is the only one whose violation cannot be\ndetected downstream.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/observation-before-action"
    },
    {
      "kind": "document",
      "slug": "protocol-aegis",
      "title": "Aegis",
      "summary": "Containment. The first protocol that acts — isolating and throttling to stop a local failure from becoming a systemic one.",
      "body": "*[VECTOR archive](/vector) › [Governance stack](/vector/the-protocol-stack) › Aegis*\n\nMost catastrophic failures in distributed systems are not caused by the component that\nfailed first. They are caused by the redistribution of its work onto components that\nwere already near capacity. Aegis exists to make a local failure stay local.\n\n## Historical context\n\nDetection and classification produce an accurate description of a system getting\nworse, and nothing else. Aegis is the point at which the chain stops observing and\nstarts acting, and it is deliberately the first protocol with that authority — because\nthe earliest possible intervention is also the cheapest and the most reversible.\n\n## Responsibilities\n\nAegis owns isolation, throttling, and the global circuit breakers. It decides which\nnodes stop receiving work and which have their intake limited.\n\nAegis never owns repair or recovery. It removes a failing component from the path; it\ndoes not attempt to fix it — that is [Mender](/vector/protocol-mender) — and it does\nnot bring it back — that is [Phoenix](/vector/protocol-phoenix). Nor does it\nredistribute the work it has displaced, which belongs to [Atlas](/vector/protocol-atlas).\n\nThe narrowness is the point. Containment that also repairs is containment that can be\ndelayed by a repair attempt.\n\n## Inputs\n\nThe stress state from [Serpentine](/vector/protocol-serpentine), and per-node error\nrates and load figures originating with Sentinel.\n\n## Outputs\n\nIsolation and throttling decisions, circuit breaker state, and — for every action — a\nlog entry carrying explicit reasoning.\n\nThe reasoning is not incidental. An isolation decision without a recorded rationale is\nindistinguishable after the fact from an arbitrary one, and the operator reviewing an\nincident needs to know why a node was removed, not merely that it was. This is\n[explainability before automation](/vector/explainability-before-automation) applied at\nthe level of a single protocol.\n\n## Internal behaviour\n\nThree mechanisms of increasing severity.\n\n**Isolation.** Nodes exceeding a 20% error rate are removed from the working set. The\nthreshold is low deliberately: a node failing one request in five is not marginal, it\nis broken, and the cost of removing a functioning node is capacity while the cost of\nretaining a broken one is propagated errors.\n\n**Throttling.** Nodes above 95% load have their intake limited rather than being\nremoved. A heavily loaded node is not faulty; it is oversubscribed, and removing it\ntransfers the oversubscription elsewhere. Throttling is the gentler instrument for the\ncase where the node is still doing useful work.\n\n**Global circuit breakers.** Under catastrophic state, breakers activate system-wide.\nThis is no longer per-node judgement — it is the recognition that individual decisions\nhave stopped being sufficient.\n\n## Interactions\n\n```\nSerpentine → Aegis → Atlas → Mender → Phoenix\n```\n\nAegis runs before Atlas, and the order is load-bearing: containment precedes\nrebalancing, so that work is never redistributed onto a node that should already have\nbeen isolated. Reversing them would move traffic onto failing capacity.\n\n## Failure modes\n\n**Over-isolation.** Aggressive removal reduces capacity, which raises load on what\nremains, which raises error rates, which triggers further isolation. This is the\ncharacteristic cascade of a containment layer, and it is caused by the protection\nrather than the fault. Throttling exists partly to provide a response that does not\nreduce capacity.\n\n**Under-isolation.** A node failing in a way that does not raise its error rate above\nthreshold — returning wrong answers rather than errors — is not caught at all. Aegis\nobserves failure, not correctness, and a component that fails silently is invisible to\nit. This is a genuine gap and not one the current design addresses.\n\n**Isolating the wrong node.** Error rates attribute failure to where it surfaced, which\nis not always where it originated. A healthy node depending on a failing one exhibits\nerrors of its own.\n\n## Design tradeoffs\n\n**Fixed thresholds across heterogeneous components.** 20% and 95% are applied\nuniformly. A component whose normal error rate is legitimately higher is isolated\nunnecessarily; one whose failure appears well below the threshold is not caught.\nPer-component thresholds would fit better and would multiply the configuration surface\nthat has to be reasoned about during an incident.\n\n**Availability is spent to buy containment.** Every isolation reduces capacity, and\nthe protocol accepts degraded throughput to prevent propagated failure. Under\nsufficiently broad degradation this is the wrong trade, and Aegis has no mechanism for\nrecognising that it has isolated too much.\n\n## Future evolution\n\nThe identified gap is silent failure — detecting components that are wrong rather than\nerroring. That requires correctness signals Aegis does not currently receive, and\nplausibly belongs upstream in detection rather than here.\n\nDependency-aware isolation, so that attribution follows the call graph rather than the\nsurface where errors appeared, is the second. Both are inference about direction, not\nplanned work.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/protocol-aegis"
    },
    {
      "kind": "document",
      "slug": "protocol-atlas",
      "title": "Atlas",
      "summary": "Rebalancing. Redistributes load across healthy capacity, and knows the difference between a distribution problem and a capacity problem.",
      "body": "*[VECTOR archive](/vector) › [Governance stack](/vector/the-protocol-stack) › Atlas*\n\nContainment removes capacity from a system that was already under pressure. Something\nhas to decide where the displaced work goes, and the naive answer — spread it evenly —\nis how a contained failure becomes a general one.\n\n## Historical context\n\nAtlas exists because [Aegis](/vector/protocol-aegis) creates a problem it deliberately\ndoes not solve. Isolation is only safe if the work the isolated node was carrying lands\nsomewhere that can absorb it. Without a rebalancing stage, containment displaces load\nblindly, and the most common outcome is a second isolation shortly afterwards.\n\n## Responsibilities\n\nAtlas owns the distribution of work across available capacity, and the judgement of\nwhen redistribution is no longer sufficient.\n\nAtlas never owns capacity itself. It does not create nodes, start containers or scale\nanything — it issues a recommendation that an autoscaler act, and the distinction is\ndeliberate. Provisioning has costs and consequences outside the boundary of this\nsystem, and a protocol operating inside a real-time loop is the wrong place to commit\nto them.\n\nNor does Atlas decide which nodes are healthy. That determination arrives from Aegis\nand Sentinel; Atlas distributes across whatever it is told remains.\n\n## Inputs\n\nPer-node load figures, the set of nodes still in service after containment, and the\nstress state from [Serpentine](/vector/protocol-serpentine).\n\n## Outputs\n\nLoad redistribution decisions, and an autoscaler activation recommendation when\nredistribution cannot reach an acceptable state.\n\n## Internal behaviour\n\nNodes above 80% load shed work to nodes below 50%, targeting 65% after the shift.\n\nThree numbers, and the gap between them is the mechanism. Moving from 80% to a target\nof 65% leaves headroom rather than levelling everything to an identical figure —\nbecause a perfectly balanced system at high utilisation has no capacity to absorb the\nnext event, and the next event is what the whole chain exists for.\n\nThe 50% ceiling on recipients prevents the obvious pathology: shifting work onto a node\nthat was itself approaching the shedding threshold, producing a cascade of\nredistribution that does no useful work and consumes budget while conditions\ndeteriorate.\n\nWhen no recipient satisfies the constraint, Atlas does not relax it. It recommends\nautoscaling. **The system is explicit that this is a capacity problem rather than a\ndistribution problem**, instead of pretending it solved something it did not — which is\nthe behaviour a rebalancer with no floor would exhibit.\n\n## Interactions\n\n```\nAegis → Atlas → Mender\n```\n\nAtlas runs after containment and before repair. It works with the node set Aegis has\nleft it, and its output changes the load figures that Sentinel will observe on the\nfollowing cycle — which makes it one of the few protocols whose action is visible to\nthe detection stage as a change in the system it is watching.\n\n## Failure modes\n\n**No valid recipient.** Handled by design: the autoscaler recommendation is the defined\noutcome rather than an error. Atlas has done its job by correctly identifying that it\ncannot help.\n\n**Redistribution thrash.** Work moved to a node that subsequently crosses the shedding\nthreshold moves again. The 50%/65% gap makes this unlikely rather than impossible, and\na system near saturation across all nodes is where it becomes plausible — which is\nexactly when the wasted movement is least affordable.\n\n**Stale load figures.** Redistribution decided from measurements that have aged is\nredistribution against a system that no longer exists. The tighter Sentinel's budget,\nthe fresher these are; this is one of the less visible benefits of the latency work.\n\n## Design tradeoffs\n\n**Deliberate under-utilisation.** Targeting 65% rather than balancing to a uniform\nfigure leaves capacity unused during normal operation. That is the cost of having\nsomewhere to put the next surge, and it is paid continuously for a benefit that is only\nvisible occasionally — the kind of tradeoff that is easy to argue away and expensive\nto have argued away.\n\n**Load is a proxy.** Utilisation percentages do not distinguish a node doing hard work\nefficiently from one struggling with easy work. Atlas balances the number, not the\nreality behind it.\n\n**Recommendation rather than action.** Deferring provisioning means the system can\nidentify that it needs more capacity and be unable to obtain it. Accepted: a real-time\nloop should not be committing to resources whose consequences extend well beyond it.\n\n## Future evolution\n\nCost-aware placement — accounting for what a particular unit of work actually demands,\nrather than treating load as fungible — is the natural direction, and would need a\nricher signal than utilisation. Predictive rebalancing, shifting work ahead of a\nforecast surge rather than in response to a measured one, would connect Atlas to the\n[intelligence layer](/vector/the-intelligence-layer), which it currently does not touch.\nBoth are inference.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/protocol-atlas"
    },
    {
      "kind": "document",
      "slug": "protocol-hyperion",
      "title": "Hyperion",
      "summary": "Emergency shutdown. Fires within one second, halts everything, and is deliberately the least sophisticated component in the system.",
      "body": "*[VECTOR archive](/vector) › [Governance stack](/vector/the-protocol-stack) › Hyperion*\n\nEvery preceding protocol assumes the system can be brought back to a working state.\nHyperion is what exists for the case where that assumption has failed, and its entire\ndesign follows from a single observation: **complexity in an emergency mechanism is\nadditional surface for the emergency to have already compromised.**\n\n## Historical context\n\nThe chain from detection through recovery handles degradation. It presumes there is\ntime to diagnose, contain, repair and restore. Some failures do not offer that, and a\nsystem whose only responses take seconds to reason through has no answer for them.\n\nHyperion is the answer, and it was built to be unlike everything above it.\n\n## Responsibilities\n\nHyperion owns exactly one action: halting the system, within one second, when\nconditions require it.\n\nHyperion never owns diagnosis, containment, repair, recovery, or any judgement about\nwhether halting was correct. It does not attempt partial shutdown, does not preserve\noperations it might consider safe, and does not reason about consequences. Each of\nthose would be another decision capable of being wrong at the exact moment nothing else\nis working.\n\n## Inputs\n\nThe conditions that trigger it. Deliberately few — every additional input is another\ndependency that must be functioning for the emergency stop to work, which inverts the\npurpose.\n\n## Outputs\n\nA halted system. There is no output stream, because there is nothing downstream: this\nis the end of automated escalation.\n\n## Internal behaviour\n\nFires within one second. Halts everything. No complex logic. No second-guessing.\nIrreversible.\n\nThe specification is that short because the mechanism is. The one-second bound is the\nbinding requirement — an emergency stop that takes ten seconds to decide is not an\nemergency stop — and it constrains everything else: there is no time for consultation,\nno time for graceful sequencing, no time for a model to be asked anything.\n\nIrreversibility is a feature. A shutdown that can undo itself is a shutdown that can be\nundone by whatever caused the emergency.\n\n## Interactions\n\n```\n(any point in the chain) → Hyperion → [halted]\n```\n\nHyperion is not a sequential stage. It is reachable from the escalation chain when\nconditions warrant it, and it does not wait for the ordinary path to complete — waiting\nfor containment and repair to be attempted is exactly what it exists to bypass.\n\nIt is the last protocol the system can invoke on its own. Beyond it is\n[Terminus](/vector/protocol-terminus), which the system may not reach.\n\n## Failure modes\n\n**Hyperion fails to fire.** The worst failure in the architecture: the last automated\nprotection is absent under exactly the conditions it exists for. The mitigation is\nsimplicity — fewer moving parts, fewer dependencies, less to have been compromised.\nThis is why sophistication here is a defect rather than an improvement.\n\n**Hyperion fires unnecessarily.** An expensive false positive: a functioning system is\nhalted. Recovery is possible but not automatic, and the cost is real. Accepted, because\nthe asymmetry is severe — an unnecessary halt is expensive, and a missing halt may not\nbe recoverable at all.\n\n**Hyperion fires too slowly.** Behaviour outside the one-second bound is undefined in\nthe sense that matters: the guarantee the rest of the system relies on has not held.\n\n## Design tradeoffs\n\n**All-or-nothing.** Halting everything is blunt. A more selective shutdown would\npreserve capability and would require judgement about which parts are safe — a judgement\nthat must be correct while the system is in a state nobody anticipated. The blunt\ninstrument is chosen precisely because it needs to be right about nothing.\n\n**No graceful degradation.** Unlike every other protocol, Hyperion has no partial\nbehaviour. It fires or it does not.\n\n**Simplicity over capability, permanently.** This is the component that must never\naccumulate features. Every future improvement to it should be regarded with suspicion,\nand that instruction is part of the specification rather than an aside.\n\n## Future evolution\n\nThe honest answer is that Hyperion should evolve as little as possible. The pressure to\nadd selectivity, preserve subsystems, or make the trigger cleverer will recur, and each\nincrement trades a reliable guarantee for a conditional one.\n\nWhere work does belong: verifying that the guarantee holds. Testing that it fires within\nits bound under adversarial conditions is worth considerably more than any feature.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/protocol-hyperion"
    },
    {
      "kind": "document",
      "slug": "protocol-mender",
      "title": "Mender",
      "summary": "Repair. The only protocol in the governance stack that uses a learned model — and the one most carefully constrained, by a budget rather than a rule.",
      "body": "*[VECTOR archive](/vector) › [Governance stack](/vector/the-protocol-stack) › Mender*\n\nContainment and rebalancing keep a degraded system running. Neither makes it well.\nMender is the protocol that attempts to repair what failed, and it is the point in\nthe governance stack where the system does something genuinely difficult to predict.\n\n## Historical context\n\nThe protocols before Mender are deterministic: thresholds, comparisons, defined\nresponses. Repair resists that treatment, because diagnosing why a component failed\nand choosing among possible remedies is a judgement rather than a lookup.\n\nThat makes Mender the first component where a learned model was admitted into the\ngovernance chain — under [governance before intelligence](/vector/governance-before-intelligence),\nthe surrounding structure had to exist first, and it did.\n\n## Responsibilities\n\nMender owns incident diagnosis and the ranking of repair strategies by risk.\n\nMender never owns the decision to restore service. A repaired component is not\nreturned to the working set by Mender; that is [Phoenix](/vector/protocol-phoenix).\nIt also does not own its own restraint — the credit budget that governs how often it\nmay act is a property of the protocol, not a judgement it makes.\n\n## Inputs\n\nIncident context: what failed, the anomalies preceding it, the containment actions\nalready taken, and the current stress state. Its own credit balance is an input to\nwhether it acts at all.\n\n## Outputs\n\nA diagnosis, and a set of repair strategies ranked by risk. Consumed downstream and by\nthe audit record — the ranking is retained, not only the strategy chosen, because an\nincident review needs to know what the alternatives were.\n\n## Internal behaviour\n\nA dual-head neural network. One head diagnoses the incident; the other ranks candidate\nrepair strategies by risk. Sharing a representation between the two reflects that the\nappropriate remedy depends on the nature of the failure, and the heads are separate\nbecause the outputs answer different questions.\n\nThe more consequential mechanism is not the network. It is the **adaptive credit\nsystem** that governs it: credits regenerate at 0.02 per second to a maximum of 10, and\nacting consumes them.\n\nThis makes aggression a budget. A system experiencing repeated failures depletes its\ncredits and becomes progressively more conservative — not because a rule detected a\nrepair loop, but because the resource that repair requires has run out. Thrashing is\nprevented structurally rather than by a heuristic that has to recognise thrashing while\nit is happening.\n\nThe regeneration rate is the design decision worth noticing. At 0.02 per second, a full\nbudget takes over eight minutes to accumulate. Repair is explicitly not something the\nsystem may do continuously.\n\n## Interactions\n\n```\nAtlas → Mender → Phoenix\n```\n\nMender runs after containment and rebalancing have stabilised the system, and before\nrecovery. The ordering means repair is attempted on a system that is no longer\nactively degrading — diagnosis against a moving target is diagnosis of the movement.\n\nIt is gated by stress state and by credit balance, and its behaviour is further\nconstrained when [the authority layer](/vector/the-authority-layer) is active.\n\n## Failure modes\n\n**Misdiagnosis.** A wrong diagnosis produces a repair aimed at the wrong problem,\nconsuming credits and changing system state without improving it. Risk ranking is the\npartial defence — preferring low-risk strategies means a wrong diagnosis is less likely\nto be destructive — but a confident wrong answer remains possible and is the residual\nrisk of admitting a model here at all.\n\n**Credit exhaustion during a genuine emergency.** The budget does not distinguish\nwasted attempts from necessary ones. A system that has legitimately needed many repairs\narrives at a real incident with nothing left. This is a deliberate acceptance: the\nalternative is an unbounded repair loop, which is worse.\n\n**Repair that worsens conditions.** Mitigated by the ranking and the budget, and not\neliminated. Escalation past Mender to shutdown exists for the case where it is not\ncontained.\n\n## Design tradeoffs\n\n**A learned model inside the governance stack.** Everything before Mender can be\nreasoned about exhaustively. This cannot. The justification is that repair genuinely\nrequires judgement — but it means the stack contains one component whose behaviour\nunder unprecedented conditions is not fully predictable, and that should be stated\nplainly rather than qualified away.\n\n**A budget rather than a rule.** A credit system is cruder than logic that detects\nrepair loops directly. It is also far harder to defeat: a rule has to recognise the\npathology, and a budget does not have to recognise anything.\n\n## Future evolution\n\nThe most valuable addition would be outcome feedback — closing the loop so that whether\na repair actually worked informs later ranking. Whether the current design does this is\nnot something the available sources state, and it should not be assumed.\n\nA credit rate that varies with stress state, so that a system in genuine crisis is not\nrationed identically to one in routine difficulty, is the other direction. Inference,\nnot planned work.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [
        {
          "revisedAt": "2026-08-07T00:00:00.000Z",
          "summary": "Renamed. The protocol repairs things; \"Punisher\" describes the opposite of what it does, and it was the most conspicuous of four Marvel names in a stack of eight. Engineers notice naming patterns, and the pattern was costing more than the name was worth.",
          "previous": "Punisher",
          "replacement": "Mender"
        }
      ],
      "url": "https://vihaanvaghela.com/vector/protocol-mender"
    },
    {
      "kind": "document",
      "slug": "protocol-phoenix",
      "title": "Phoenix",
      "summary": "Recovery. Brings the system back in dependency order, and deliberately restarts more than strictly necessary.",
      "body": "*[VECTOR archive](/vector) › [Governance stack](/vector/the-protocol-stack) › Phoenix*\n\nRestoring a distributed system is not the reverse of failing. Components come back in\nan order that matters, carrying state that may or may not be valid, into a system whose\nshape has changed while they were absent. Phoenix owns that process.\n\n## Historical context\n\nRepair without recovery leaves a system that is technically fixed and not actually\nserving. Something has to decide when a repaired component rejoins, in what order, and\nwith what state — and the naive approach, restarting everything simultaneously, tends\nto produce a thundering-herd effect against dependencies that have not themselves\nrecovered.\n\n## Responsibilities\n\nPhoenix owns restart sequencing, container rebuilds, and the restoration of global\nruntime state.\n\nPhoenix never owns the decision that a component is repairable — that is\n[Mender](/vector/protocol-mender) — and it does not decide whether the system is\nsafe to operate, which remains with [Serpentine](/vector/protocol-serpentine) and\n[the authority layer](/vector/the-authority-layer). Phoenix restores; it does not judge.\n\n## Inputs\n\nThe set of components requiring restoration, the dependency graph between them, per-node\nload, and the runtime state to be restored.\n\n## Outputs\n\nSequenced restarts, container rebuilds, and a restored global runtime state.\n\n## Internal behaviour\n\nThree mechanisms.\n\n**Dependency-ordered restarts.** Components come back in an order that respects what\nthey depend on. A service restarted before its dependency fails immediately and\nrequires another restart, which is both slower and noisier than waiting.\n\n**Container rebuilds above 85% load.** A heavily loaded node is rebuilt rather than\nrestarted. The distinction assumes that a node in that condition may carry accumulated\nstate contributing to its problem — a rebuild discards it, at the cost of a slower\nreturn.\n\n**Global runtime state restoration.** Recovered components rejoin a system with\ncoherent shared state rather than reconstructing it independently and disagreeing.\n\nThe governing disposition is conservatism, stated explicitly: **Phoenix performs more\nrestarts than strictly necessary, never fewer.** Restarting something that did not need\nit costs milliseconds. Failing to restart something that did costs the recovery, and\nthe failure is discovered later, in a system everyone believes has recovered.\n\n## Interactions\n\n```\nMender → Phoenix → (system returns to normal operation)\n```\n\nPhoenix is the last protocol in the ordinary path. Beyond it the chain contains only\n[Hyperion](/vector/protocol-hyperion) and [Terminus](/vector/protocol-terminus), which\nare not part of normal operation.\n\nIts output is observed by Sentinel on subsequent cycles, which is how recovery is\nconfirmed: the system does not assert that it has recovered, it observes that it has.\n\n## Failure modes\n\n**Restart loop.** A component that fails on restart is restarted again. Without a\nbound this consumes the recovery indefinitely; whether Phoenix bounds its own attempts\nis not stated in the available sources and should not be assumed.\n\n**Restored state that is invalid.** State captured during degradation may encode the\ndegradation. Restoring it returns the system to the condition it was recovering from —\nthe failure mode a rebuild is designed to escape, which is why the load threshold\nexists.\n\n**Recovery under continuing stress.** Bringing components back into a system still\nunder pressure adds load to something already struggling. Sequencing limits the rate;\nit does not remove the risk.\n\n## Design tradeoffs\n\n**Deliberate over-restarting.** Recovery is slower than the minimum, always. The\nalternative — restarting exactly what is needed — requires knowing exactly what is\nneeded, and being wrong is discovered late and cheaply hidden. This is a preference for\na small known cost over a rare large one.\n\n**Rebuild versus restart is decided by load.** A single proxy for a judgement about\naccumulated state. It is a reasonable heuristic and it is a heuristic.\n\n## Future evolution\n\nVerified recovery — confirming a component is actually healthy before proceeding to the\nnext in the sequence, rather than inferring it from a successful restart — is the\nclearest direction. Partial recovery, restoring the most critical path first and the\nremainder afterwards, is the second. Both are inference.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/protocol-phoenix"
    },
    {
      "kind": "document",
      "slug": "protocol-sentinel",
      "title": "Sentinel",
      "summary": "Detection. The first stage of the escalation chain, the only component that observes the system directly, and the one that nearly made the system unusable.",
      "body": "*[VECTOR archive](/vector) › [Governance stack](/vector/the-protocol-stack) › Sentinel*\n\nA system cannot govern what it cannot see. Sentinel is the component that sees, and\neverything downstream of it is acting on a description of the system rather than the\nsystem itself — which makes the fidelity of that description the upper bound on the\nquality of every decision that follows.\n\n## Historical context\n\nSentinel was the first protocol built, and for a period it was the only one. A system\nthat merely reported anomalies to a human was already more useful than one that did\nnot, and everything else in the escalation chain was added because reporting turned\nout to be insufficient at machine speed.\n\nIt is also the component that nearly ended the project. Its original implementation\nperformed unbounded statistical work on every cycle and consumed roughly 97% of\nruntime, holding p95 latency near 100 ms — while being the last component anyone\nsuspected, because it was the simplest and the most trusted. The full account is in\n[the latency crisis](/vector/the-latency-crisis), and the constraint it produced —\nthat every stage carries a hard time budget — now applies system-wide.\n\n## Responsibilities\n\nSentinel owns the collection of runtime metrics, the detection of statistical\nanomalies within them, and the categorisation and severity scoring of what it finds.\n\nSentinel never owns interpretation. It does not decide whether the system is in\ntrouble — that is Serpentine's judgement, formed from Sentinel's output. It does not\nact, contain, throttle or repair. The separation is deliberate and follows\n[the boundary between seeing and deciding](/vector/the-boundary-between-seeing-and-deciding):\na detector that also decides inherits every one of its own weaknesses into the\ndecision, silently.\n\n## Inputs\n\nRuntime telemetry, sampled every second: latency, CPU, memory, GPU, queue depth and\nthroughput. Sentinel consumes no output from any other protocol, which makes it the\nonly component in the chain with no upstream dependency — and the reason it can still\nfunction when everything above it has failed.\n\n## Outputs\n\nTwo streams, deliberately separated by priority.\n\nRaw metrics travel at medium priority: they are the ordinary substrate other\ncomponents reason from. Detected anomalies escalate at high priority immediately.\n\nThe priority split matters more than it appears. A bus that treats \"something is\nwrong\" as ordinary traffic delivers that message behind whatever queue had already\nformed — which is longest precisely when something is wrong.\n\n## Internal behaviour\n\nDetection is statistical rather than threshold-based: z-score analysis against a\nrolling history of twenty samples. A fixed threshold answers \"is this value large\",\nwhich requires knowing in advance what large means for every metric under every\ncondition. A rolling comparison answers \"is this value unusual for this system right\nnow\", which is the question that generalises.\n\nThe window is short by design. Twenty samples adapts quickly enough to follow genuine\nregime changes rather than alarming through all of them, at the cost of adapting to a\ndegradation that arrives slowly enough — a tradeoff discussed below.\n\nSince the latency work, statistics are maintained incrementally at constant cost per\nsample, heavier analysis is displaced into bounded background tasks, and the whole\nstage runs under a 5 ms hard cap with early exit when risk or latency indicators\nexceed thresholds.\n\n## Interactions\n\n```\ntelemetry → Sentinel → Serpentine → (rest of the chain)\n```\n\nSentinel executes first, every cycle, unconditionally. Its only consumer in the\ngovernance chain is [Serpentine](/vector/protocol-serpentine), which synthesises its\noutput into a single stress state. Its anomaly stream is also what the intelligence\nlayer's trigger conditions are evaluated against — see\n[the intelligence layer](/vector/the-intelligence-layer).\n\nNothing calls Sentinel. It is not a service that answers questions; it is a stage that\nruns.\n\n## Failure modes\n\n**Sentinel stops producing.** The chain has no input. Every downstream protocol is\nreasoning from a stale description of a system that is still changing, and — worse —\nnothing downstream can distinguish \"no anomalies\" from \"no observations\". This is the\nmost dangerous failure in the stack, because it is silent and it looks like health.\nThe mitigation is that absence of output is itself an observable condition rather than\nan inferred one.\n\n**Sentinel exceeds its budget.** Bounded by the 5 ms cap: the stage exits with what it\nhas rather than running long. Partial detection is a degradation. An unbounded head of\nthe pipeline is a system-wide failure, which is exactly what the latency crisis was.\n\n**Sentinel detects spuriously.** Excess anomalies propagate as elevated stress, which\nnarrows autonomy unnecessarily. This is the correct direction to fail in — it costs\ncapability rather than safety — but sustained false positives erode the credibility of\nthe whole chain, which is a slower and more serious harm.\n\n## Design tradeoffs\n\n**The rolling window is short.** Twenty samples means a fault that degrades slowly\nenough can become the new baseline without ever registering as anomalous. Detecting\nthat class of drift needs a longer horizon, and a longer horizon adapts too slowly to\nlegitimate regime change. There is no window length that solves both; this one favours\nresponsiveness, and the gap is real.\n\n**Statistical detection has no semantics.** Sentinel knows a value is unusual. It does\nnot know whether unusual is bad. That judgement is deferred to Serpentine, which keeps\nthis component simple and means detection carries no domain knowledge at all.\n\n**The time cap can hide work.** A stage that exits early under load is doing less\ndetection at precisely the moment there is more to detect. Accepted deliberately: late\ndetection in a real-time system is indistinguishable from no detection, and the\nunbounded alternative was measured and was worse.\n\n## Future evolution\n\nThe clearest gap is adversarial input. Sentinel currently assumes its telemetry is\nhonest but noisy. A sensor that is wrong deliberately — reporting plausible values\nthat are false — defeats statistical detection entirely, because the readings are\ndesigned to be unremarkable. An adversarial robustness layer is identified as future\nwork in [the v1 changelog](/vector/changelog-v1) and is not implemented.\n\nThe second is multi-horizon detection: maintaining several windows simultaneously so\nthat slow drift and fast spikes are both visible without one window having to serve\nboth. This is inference about a natural direction rather than a planned change.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/protocol-sentinel"
    },
    {
      "kind": "document",
      "slug": "protocol-serpentine",
      "title": "Serpentine",
      "summary": "Classification. Turns many independent signals into one judgement about system state, so that everything downstream reasons from a shared answer.",
      "body": "*[VECTOR archive](/vector) › [Governance stack](/vector/the-protocol-stack) › Serpentine*\n\nA stream of anomalies is not a description of a system's condition. Six components\neach reporting mild irregularity may indicate nothing, or may indicate the early\nstage of a cascade, and no individual signal contains the difference. Serpentine\nexists to convert many local observations into one global judgement.\n\n## Historical context\n\nThe need became apparent once more than one protocol consumed Sentinel's output.\nEach downstream component was applying its own thresholds to the same anomaly stream,\nwhich meant they could disagree about whether the system was in trouble — and a\ncontainment layer that believes conditions are critical while a rebalancing layer\nbelieves they are stable produces incoherent behaviour that is very difficult to\ndiagnose after the fact.\n\nCentralising the judgement is not an efficiency measure. It is what makes the rest of\nthe chain reason from a single shared premise.\n\n## Responsibilities\n\nSerpentine owns the definition of system stress: the indices it is computed from, the\nthresholds at which one level becomes another, and the single classification that is\nthe authoritative answer to \"how is the system doing\".\n\nSerpentine never owns response. It classifies and stops. Nothing about containment,\nrebalancing, repair or shutdown belongs here, and the temptation to let the component\nthat knows the state also act on it is exactly the coupling the architecture refuses.\n\n## Inputs\n\nEverything [Sentinel](/vector/protocol-sentinel) produces: raw metrics and categorised,\nseverity-scored anomalies.\n\n## Outputs\n\nOne stress state, from five levels: **stable, moderate, critical, failure,\ncatastrophic**. Every protocol downstream reads this value.\n\nThe states are ordered and the ordering is meaningful — this is a severity scale, not\na set of categories — which is what allows downstream components to express their\nbehaviour as thresholds against it rather than as a table of cases.\n\n## Internal behaviour\n\nThe classification is derived from three indices rather than directly from raw\nmetrics: **pressure**, **congestion** and **instability**.\n\nThe intermediate layer is doing real work. It separates three genuinely different\nways a system can be in difficulty — pressure is load approaching capacity, congestion\nis work accumulating faster than it drains, instability is behaviour that will not\nsettle. A single scalar cannot distinguish them, and the appropriate response differs:\npressure suggests rebalancing, instability suggests narrowing autonomy.\n\nThresholds are deliberately severe. Catastrophic requires 92% pressure or 95%\ninstability, and not lower. A classification that fires early is a classification\nnobody believes, and an escalation chain nobody believes has become decoration —\noperators route around a system that cries wolf, which removes the protection while\nleaving the appearance of it.\n\n## Interactions\n\n```\nSentinel → Serpentine → Aegis, Atlas, Mender, Phoenix, Hyperion\n                     ↘ intelligence layer gating\n```\n\nSerpentine executes second, immediately after Sentinel, before any intelligence runs.\nThat ordering is the architectural claim: **the system establishes how stressed it is\nbefore it decides how ambitious to be.** Its output gates the intelligence layer, and\nis one of the signals feeding [the authority layer](/vector/the-authority-layer).\n\n## Failure modes\n\n**Misclassification downward** — reporting stable while conditions deteriorate — is\nthe serious one. Every protective response in the chain is gated on this value, so an\nunderstated classification disables all of them simultaneously. The severe thresholds\ntrade against this: they make catastrophic harder to reach, which is the correct\nchoice for credibility and the wrong one for this failure. The tension is real and\nunresolved.\n\n**Misclassification upward** costs capability. Containment engages unnecessarily,\nintelligence is gated off, the system is more conservative than conditions warrant.\nRecoverable, and the direction one prefers to fail in.\n\n**Oscillation across a threshold** produces repeated engagement and disengagement of\ndownstream responses, which is worse than sitting on either side of it. Sustained-signal\nrequirements elsewhere in the chain — notably in the authority layer — exist partly to damp this.\n\n## Design tradeoffs\n\n**One global state for a system that may be locally unwell.** A single classification\ncannot express \"node seven is failing and everything else is fine\". Aegis operates\nper-node for exactly this reason, but the global state remains a summary, and summaries\nlose things.\n\n**Fixed thresholds in an adaptive system.** The levels are constants, not learned. That\nis consistent with [governance before intelligence](/vector/governance-before-intelligence)\n— the boundary must not move in response to pressure applied to it — and it means the\nthresholds cannot adapt to a system whose normal operating range has legitimately\nchanged.\n\n## Future evolution\n\nThe natural extension is regional classification: stress states per subsystem beneath\nthe global one, so containment can be scoped without losing the shared premise. This\nis inference rather than a stated plan, and it carries an obvious hazard — several\nstress states can disagree, which reintroduces the incoherence Serpentine was built to\nremove.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/protocol-serpentine"
    },
    {
      "kind": "document",
      "slug": "protocol-terminus",
      "title": "Terminus",
      "summary": "Permanent termination, reachable only by a human. The system cannot invoke the end of its own escalation chain — and that is the most important property in the design.",
      "body": "*[VECTOR archive](/vector) › [Governance stack](/vector/the-protocol-stack) › Terminus*\n\nEvery protocol before this one is the system managing itself. Terminus is the point at\nwhich it is not permitted to, and it is the structural expression of\n[the human remains the operator](/vector/the-human-remains-the-operator).\n\n## Historical context\n\nA system that can detect, contain, repair, recover and halt is a system with a complete\nautomated response to its own failure. That completeness is the problem. If every step\nof the chain is reachable by the system, then the system's authority has no boundary —\nit merely has a longest path.\n\nTerminus was built to be the step that is not reachable. Its purpose is not to end the\nsystem; it is to establish that ending the system is something only a person can do.\n\n## Responsibilities\n\nTerminus owns permanent termination, the authorisation required to invoke it, and the\npreservation of system state before anything ends.\n\nTerminus never owns any automated decision. It has no trigger conditions, no thresholds,\nand no relationship to stress state. Nothing it does is a response to a measurement.\n\n## Inputs\n\nA human being, physically present at a console, with multi-factor authorisation.\n\nThat is the complete list, and each element is doing work. **Physically present**\ndefeats remote invocation by anything that has obtained credentials. **Multi-factor**\ndefeats a single compromised factor. **Human** is enforced rather than assumed:\norchestrators, AI agents and internal services are explicitly blocked from invoking it.\n\nThat blocklist is the specification. A termination path that is merely undocumented is\nreachable by anything that discovers it, and the whole property collapses.\n\n## Outputs\n\nA terminated system, and an encrypted full-state archive written before termination\ncompletes.\n\nThe archive is not incidental. Termination is invoked when something has gone wrong\nenough that a person has decided the system should not continue, which is precisely when\nthe evidence is most needed and least likely to survive. Encryption reflects that a full\nsystem state contains everything the system knew.\n\n## Internal behaviour\n\nAuthorisation is verified. State is captured and encrypted. The system ends.\n\nRecovery afterwards is possible — but deliberately, and only by a person. There is no\nautomatic restart, no watchdog that notices the system is gone and brings it back. A\ntermination that something else can reverse is not a termination.\n\n## Interactions\n\n```\nSentinel → Serpentine → Aegis → Atlas → Mender → Phoenix → Hyperion  │  Terminus\n└──────────────── reachable by the system ──────────────────────────┘  └─ human only ─┘\n```\n\nTerminus sits outside the chain rather than at the end of it. Nothing calls it. It has\nno upstream protocol, consumes no telemetry, and cannot be escalated to.\n\n[Hyperion](/vector/protocol-hyperion) is the last step the system can take on its own,\nand the boundary between the two is the boundary of the system's authority.\n\n## Failure modes\n\n**Terminus is unavailable when needed.** The operator cannot end a system that should\nbe ended. This is the failure that matters, and it is why availability of the\ntermination path is a property to be tested rather than assumed — a path exercised only\nin circumstances nobody wants to rehearse is a path nobody knows works.\n\n**Terminus is invoked when it should not have been.** Expensive and, importantly,\nlegitimate: the authority to end the system includes the authority to be wrong about it.\nThe archive is what makes the decision reviewable afterwards.\n\n**The blocklist is incomplete.** A path by which an automated component reaches Terminus\ndefeats the entire property, silently, and would probably only be discovered by being\nused. This is the failure mode to audit for rather than reason about.\n\n## Design tradeoffs\n\n**Human presence in an emergency.** Requiring physical presence and multi-factor\nauthorisation makes termination slow, and slow is a real cost when it is needed. That\ncost is why [Hyperion](/vector/protocol-hyperion) exists: the fast stop is automated and\nreversible, the permanent stop is deliberate and not.\n\n**Deliberate incompleteness of automation.** The system is built with a capability it\ncannot use. That is not an oversight to be resolved in a later version — it is the\ndesign, and any future change that makes Terminus reachable by an automated component\nshould be understood as removing the property rather than improving the protocol.\n\n## Future evolution\n\nTerminus should change less than anything else in this archive. The valuable work is\nassurance rather than capability: verifying the blocklist genuinely holds, that the\narchive is complete and restorable, and that the authorisation path works when it is\nneeded rather than only when it is convenient to test.\n\nIf a future version of VECTOR operates across multiple sites, the question of what\n\"physically present\" means becomes non-trivial and is not currently answered.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/protocol-terminus"
    },
    {
      "kind": "document",
      "slug": "provenance-as-a-schema",
      "title": "Provenance belongs in the schema, not the README",
      "summary": "Why this site stores essays and VECTOR documents in two different tables, and why the publish path refuses rather than warns.",
      "body": "A note about the archive you are currently reading, rather than about VECTOR.\n\nThis site stores personal essays and VECTOR documents in two separate tables. They\nhold what looks like the same thing — a slug, a title, a body of markdown, a\npublication status — and merging them would remove a model, a set of queries and a\nstudio screen.\n\nThey are separate because they have different provenance, and provenance that lives in\na README is a claim rather than a property.\n\n## The two rules\n\n**An essay is human-written.** The `Post` model has no field recording an AI\ncontribution, and the absence is the design: there is never one to record. The moment\nthat model gains an `aiAssisted` column, the rule has changed, and the schema will say\nso before any prose does.\n\n**A VECTOR document is AI-assisted and then reviewed.** `VectorDoc` carries\n`aiAssisted`, `reviewedAt` and `reviewedBy`, because that is the actual pipeline. The\ndocumentation is drafted with assistance and then read by a person.\n\n## Why refusal rather than a warning\n\nPublishing an AI-assisted document with no `reviewedAt` is refused in code.\n\nThe alternative — a warning, a checklist item, a note in the contributing guide —\ndepends on someone caring at the exact moment they are trying to ship something. That\nis the worst possible moment to rely on. A step that nothing enforces is a step that\ngets skipped, and the first time it gets skipped is the first time the site tells a\nreader something untrue about how its own contents were made.\n\nThe honest way to say \"a human read this\" is to make it impossible to publish\notherwise.\n\n## The generalisation\n\nThis is the same argument as [Terminus](/vector/protocol-terminus). If a property\nmatters, encode it where it cannot be forgotten — in the schema, in the type system, in\nthe path that has to succeed. Documentation describes intentions. Structure enforces\nthem. The companion case, where a property is deliberately *not* enforced by structure\nand the honest thing is to say so, is\n[a console with no URL](/vector/the-console-with-no-url).\n\nEvery document in this archive seeded as an unreviewed draft, and none of them could\nreach you without a person opening them first.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/provenance-as-a-schema"
    },
    {
      "kind": "document",
      "slug": "seven-variables",
      "title": "How I reduced a complex system to seven variables",
      "summary": "Complexity is easy to add and looks like intelligence. Most of it is redundant, and reduction is what survives degradation.",
      "body": "Complexity is easy to add. Any system can be made to look sophisticated by increasing\nthe number of inputs, sensors, rules and exceptions. The result often looks\nintelligent, and the appearance is fragile — when conditions change, the system loses\ncoherence, because it never learned which signals mattered.\n\nThe problem is not that real-world systems are complex. The problem is that most of\nthat complexity is redundant.\n\n## Signals that move without meaning\n\nIn environments that evolve over time — traffic, weather, human behaviour — a great\nmany signals fluctuate without carrying causal information. Treating all variation as\nmeaningful produces systems that react constantly and understand nothing.\n\nReduction here is not simplification for convenience. It is an attempt to concentrate\nattention on what persists when noise increases and assumptions fail.\n\n## The constraint\n\nFrom the beginning I imposed a rule: the system would operate on a minimal set of\nvariables that remain informative under degradation. Not driven by hardware limits or\nconvenience — a design choice.\n\nThe reasoning is about failure rather than efficiency. Systems that depend on a wide\narray of signals depend on them unevenly, and a failure in one part of the input space\ncascades unpredictably, because nobody knows which downstream behaviour was quietly\nrelying on the signal that just died.\n\nA smaller variable set means each variable is understood, its failure mode is known,\nand there is a defined answer for what the system does without it.\n\n## The test\n\nThe question for admitting a variable is not \"does this improve accuracy\". Almost\nanything improves accuracy on a benchmark. The question is: **does this remain\ninformative when conditions degrade, and do I know what the system does when it is\ngone?**\n\nMost candidate inputs fail the second half. They improve the average case and have no\ndefined behaviour in the failure case, which means adding them trades a visible\nbenchmark gain for an invisible fragility.\n\nSeven survived that test. The number is not the point — the test is.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/seven-variables"
    },
    {
      "kind": "document",
      "slug": "the-authority-layer",
      "title": "The authority layer: discipline over chaos",
      "summary": "The authority layer. When uncertainty rises, autonomy is deliberately narrowed — and the system has to earn it back in stages.",
      "body": "This is the final authority layer, and the design principle it encodes is the one I\nwould keep if I had to throw away everything else in the system:\n\n**When uncertainty rises, autonomy should narrow — not widen.**\n\nThe instinct runs the other way. A system in trouble looks like a system that needs\nmore capability brought to bear. In practice, trouble is exactly when your models are\nleast reliable, because trouble means conditions have left the distribution they were\ntrained on. The moment you most want intelligence is the moment you can least trust it.\n\n## Escalation\n\n```\nnormal → warning → critical → vihaan_activated\n```\n\nDeterministic, driven by sustained systemic risk signals rather than instantaneous\nones. A spike does not trigger the authority layer; a spike that does not resolve does.\n\n## What activation does\n\nOn activation it applies authoritative overrides:\n\n- Disables MORL and PULSAR\n- Caps oracle weight\n- Freezes the meta layer\n- Forces the baseline decision path\n\nThe system is stripped back to the deterministic controller that existed before any\nlearning was added. This is why that controller was built first, and why it is\nmaintained as a first-class component rather than left as scaffolding — it is the\nfloor the whole structure falls back to.\n\n## Earning it back\n\n```\nstabilize → monitor → gradual_reintroduction → full_restore\n```\n\nRecovery is a state machine, not a flag. Intelligence is reintroduced in stages, each\ngated on evidence of stability rather than elapsed time. The system earns back its own\nautonomy.\n\nThe failure mode this prevents is oscillation: restore everything the moment the\nmetrics look acceptable, destabilise again, activate again. Staged reintroduction with\nevidence gates is slower and it converges.\n\n## Memory\n\nIt persists memory signatures and audit logs for replay and governance learning.\nEvery decision is traceable after the fact.\n\nThat matters most for the activations that turned out to be unnecessary. A protocol\nthat narrows autonomy will sometimes narrow it wrongly, and without a replayable\nrecord there is no way to distinguish \"the authority layer saved this\" from \"the\nauthority layer panicked\" — which\nwould make the threshold impossible to tune honestly.\n\n## On the name\n\nIt is named after me, which is either the most or least defensible naming decision in\nthe project. The reasoning: this is the layer that encodes what I am willing to let\nthe system do without asking. If any component should carry the name of the person\nresponsible for it, it is that one.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [
        {
          "revisedAt": "2026-08-07T00:00:00.000Z",
          "summary": "Renamed. A core protocol inside my own system, on my own site, named after me — whatever the backronym meant to me, to a stranger it reads as self-mythologising, and it was the one name a hostile reader would screenshot. The old URL still resolves.",
          "previous": "The VIHAAN protocol",
          "replacement": "the authority layer"
        }
      ],
      "url": "https://vihaanvaghela.com/vector/the-authority-layer"
    },
    {
      "kind": "document",
      "slug": "the-boundary-between-seeing-and-deciding",
      "title": "The boundary between seeing and deciding",
      "summary": "Most failures in complex systems are not broken components. They are boundaries that were never properly drawn.",
      "body": "In complex systems, most failures do not come from individual components breaking.\nThey come from boundaries being poorly defined. When responsibilities blur, a system\nbecomes hard to reason about, harder to modify, and fragile under pressure.\n\nThe boundary that took the most work to draw properly was the line between perception\nand decision.\n\n## The tempting collapse\n\nIt is very natural to merge them. Modern models are powerful enough that you can ask a\nperception system not only to report what is present but to infer what should happen\nnext. One component, fewer interfaces, less plumbing.\n\nIn practice this produces systems that react quickly and reason poorly. The tighter the\ncoupling, the more brittle the behaviour — because the decision inherits every\nweakness of the perception, silently, with no place in the architecture where you could\nhave caught it.\n\n## What the separation buys\n\nKeeping them apart costs an interface and buys three things.\n\n**Uncertainty survives the trip.** When perception hands over a structured observation\nrather than a recommendation, it can also hand over how sure it is. A merged component\nhas no natural place to put that, so confidence gets absorbed into the output and\ndisappears. Every downstream safety gate in VECTOR depends on confidence being a\nfirst-class value, and that is only possible because the boundary exists.\n\n**Failures stay attributable.** When a decision is wrong, the separation lets you ask\nwhich half was wrong. Did it see incorrectly, or see correctly and choose badly? In a\nmerged component that question has no answer, and a question with no answer is a bug\nyou cannot fix.\n\n**Components become replaceable.** Perception can be retrained, swapped or degraded\nwithout the decision layer noticing, as long as the contract holds. This is what let\nthe system survive its own scope change.\n\n## The general rule\n\nA boundary is worth its interface cost when the two sides can fail independently. If\nperception can be wrong in ways decision-making cannot detect, they are different\ncomponents, and merging them for convenience means choosing to be unable to tell which\none failed.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-boundary-between-seeing-and-deciding"
    },
    {
      "kind": "document",
      "slug": "the-console-with-no-url",
      "title": "A console with no URL, and why that is not security",
      "summary": "The operator gateway is reached by typing a trigger into the search bar. That is a user-interface decision, and saying so out loud is what stops it becoming the plan.",
      "body": "A second note about this site.\n\nThe VECTOR operator console has no address. It is not linked, not in the sitemap, not\nin `robots.txt`. It is reached by typing a trigger string into the search bar on the\npublic site.\n\nEvery path belonging to it returns a bare 404 to an unauthenticated request —\nbyte-identical to the 404 for a path that has never existed. There is no login page to\nfind, no redirect that reveals the route exists, and no difference in timing or\nresponse body between \"you may not have this\" and \"this is not here\".\n\n## The part that matters\n\n**None of that is the security model.**\n\nEvery credential check happens server-side. The trigger string is in the JavaScript\nbundle, because it has to be — the search bar runs in the browser. Anyone willing to\nread the bundle can extract it in about a minute, and they arrive at exactly the same\nlocked door as someone who guessed: an Argon2id password, then a one-time code sent to\na mailbox they do not control, then a WebAuthn passkey bound to hardware they do not\nhave.\n\nThe obscurity removes the console from view. It removes nothing from an attacker who\nhas looked.\n\n## Why write that down\n\nBecause obscurity is load-bearing the moment nobody says it is not.\n\nThe failure mode is gradual and entirely reasonable at each step. A surface is hidden.\nBecause it is hidden, it feels lower-risk. Because it feels lower-risk, the next\ncontrol on it gets deferred — not decided against, just deferred. Repeat, and the\nhiding has quietly become the protection, and nobody ever chose that.\n\nWriting \"this is a UI property and not a control\" in the place where the decision lives\nis what stops that drift. It costs one sentence and it removes the option of pretending\nlater.\n\n## The rule\n\nHiding a thing is a fine user-interface decision. It is never a control. Both\nstatements have to be true at the same time, and the second one has to be written\nsomewhere the first one cannot outgrow it.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-console-with-no-url"
    },
    {
      "kind": "document",
      "slug": "the-data-pipeline",
      "title": "The data pipeline",
      "summary": "Multi-resolution resampling, feature engineering and walk-forward validation — and why the splitting strategy matters more than the model.",
      "body": "*[VECTOR archive](/vector) › Architecture › The data pipeline*\n\nMost of the difference between a model that works in evaluation and one that works in\nproduction is decided before any model is trained, in how the data was split. The data\npipeline is where that decision lives.\n\n## Problem\n\nTraffic is a time series, and time series break the assumption that most machine\nlearning tooling is built on. Randomly partitioning observations into training and test\nsets lets the model learn from the future and be evaluated on the past. The resulting\nscores are excellent and meaningless.\n\nThe failure is silent, and it is flattering, which is a bad combination.\n\n## Context\n\nThe environment is non-stationary — see\n[the philosophy](/vector/the-philosophy) on why variable environments are named in the\nsystem's definition. Conditions have regimes, and phenomena occur at more than one\ntimescale: a signal cycle, a rush hour and a seasonal pattern are all real, and a single\nsampling resolution captures at most one of them well.\n\n## Constraints\n\nEvaluation has to reflect deployment, which means training only on the past. Multiple\ntimescales have to be available simultaneously. And the output has to be consumable by\nthe training code without a second transformation step where assumptions can diverge.\n\n## Alternatives considered\n\n**Random splits.** Standard, convenient, and invalid for time series for the reason\nabove.\n\n**A single held-out tail.** Train on everything before a date, test after it. Honest,\nand it evaluates the model against exactly one period — so a model that happens to suit\nthat period looks better than it is.\n\n## Chosen architecture\n\n`DataPipeline` in `core/data_pipeline.py`, built around three stages.\n\n**Multi-resolution resampling** (`resample_multi_resolution`) produces several\nresolutions of the same series rather than committing to one, so phenomena at different\ntimescales remain visible.\n\n**Feature engineering** (`engineer_features`) operates per resolution and returns an\n`EngineeredDatasetBundle` — the features and the metadata describing them travel\ntogether, which is what stops a feature set and its description drifting apart.\n\n**Walk-forward splitting** (`walk_forward_splits`, producing `WalkForwardSplit`) trains\non a window and evaluates on the period immediately following it, repeatedly, advancing\nthrough the series. Every evaluation is a genuine forecast, and the model is scored\nacross many regimes rather than one.\n\nThe result is exposed as a `TimeSeriesTorchDataset` via `build_torch_dataset`, so the\nboundary between preparation and training is a typed dataset rather than a convention.\n\n## Subsystem relationships\n\nFeeds the learned components in `learning/` and `models/`. It does not feed the\ngovernance stack: the protocols operate on live telemetry through\n[the event bus](/vector/the-event-bus), not on engineered datasets. Training and\noperation are deliberately separate paths.\n\n## Data flow\n\n```\nraw series → resample_multi_resolution → {resolution: frame}\n           → engineer_features → EngineeredDatasetBundle\n           → walk_forward_splits → [WalkForwardSplit, …]\n           → build_torch_dataset → TimeSeriesTorchDataset\n```\n\n## Tradeoffs\n\n**Walk-forward is expensive.** Many train/evaluate cycles instead of one. The cost is\ncompute, and it buys an estimate that means something.\n\n**Multi-resolution multiplies the feature space,** which is in tension with the\nreduction argued for in [seven variables](/vector/seven-variables). Resolution breadth\nand variable count are separate axes, and the resolution of that tension is not visible\nfrom the pipeline alone.\n\n**Recency versus regime coverage.** A model validated across many regimes is not\nnecessarily the best model for the current one.\n\n## Failure modes\n\n**Leakage through a feature.** Walk-forward prevents leakage through splitting; a\nfeature computed over the whole series before splitting reintroduces it, and the split\nstrategy cannot detect that.\n\n**Insufficient history.** Walk-forward needs enough series to produce meaningful\nwindows; short histories yield few splits and a noisy estimate.\n\n**Regime change after training.** The pipeline measures performance across past\nregimes. It cannot say anything about a regime that has not occurred — which is why the\nruntime system has [the authority layer](/vector/the-authority-layer) rather than relying on\nvalidation.\n\n## Future evolution\n\nThe clearest gap is a feature store with explicit point-in-time semantics, which would\nmake leakage through feature computation structurally impossible rather than a\ndiscipline. Inference, not a documented plan.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-data-pipeline"
    },
    {
      "kind": "document",
      "slug": "the-event-bus",
      "title": "The event bus",
      "summary": "Why the protocols communicate through a priority queue instead of calling each other, and what that buys when the system is under load.",
      "body": "*[VECTOR archive](/vector) › Architecture › The event bus*\n\nThe governance stack is described as a chain, and a chain implies each stage calling\nthe next. It does not. The protocols never call each other; they publish and subscribe\nthrough a shared asynchronous bus, and that indirection is load-bearing.\n\n## Problem\n\nA chain of direct calls couples every stage to the stage after it. Detection has to\nknow that classification exists, classification has to know about containment, and any\nchange to the ordering is a change to every component in it.\n\nWorse for this system: a direct call inherits the callee's latency. If containment is\nslow, detection is slow, because detection is blocked inside it — and detection is the\nstage whose timeliness everything else depends on.\n\n## Context\n\nThe stack has nine protocols with a defined execution order, a hard per-cycle time\nbudget, and a requirement that \"something is wrong\" reaches its consumers faster than\nordinary telemetry does. Several components consume the same events — Sentinel's\nanomaly stream is read by classification, by the intelligence layer's trigger\nconditions, and by the audit record.\n\n## Constraints\n\nDelivery must be bounded: an unbounded queue under sustained overload converts a\nlatency problem into a memory problem, which fails later and worse. Ordering has to be\npreserved where it is meaningful. And priority must be expressible, because the whole\npoint is that alarms are not ordinary traffic.\n\n## Alternatives considered\n\n**Direct calls.** Simplest, and rejected for the coupling and latency inheritance\nabove.\n\n**A plain FIFO queue.** Decouples the stages and cannot express urgency. An anomaly\npublished behind a backlog of routine metrics is delivered after that backlog — and the\nbacklog is longest exactly when the anomaly matters most. This is the failure mode the\ndesign specifically rejects.\n\n## Chosen architecture\n\n`AsyncEventBus` in `core/event_bus.py`: an `asyncio.PriorityQueue` with a bounded\nmaximum size, topic-based `subscribe`, and events carried as a typed `Event` wrapped in\nan internal `_PrioritizedEvent` for ordering.\n\nTwo properties follow. Priority is a first-class attribute of publication rather than a\nconvention, so a high-priority anomaly overtakes queued telemetry by construction. And\nthe queue is bounded, so overload manifests as backpressure at a known limit rather\nthan as unbounded growth.\n\nTopic-based subscription is what makes one anomaly stream serve several consumers\nwithout any of them knowing the others exist.\n\n## Subsystem relationships\n\nEvery protocol in [the governance stack](/vector/the-protocol-stack) is a publisher, a\nsubscriber, or both. [Sentinel](/vector/protocol-sentinel) is the primary publisher and\nsubscribes to nothing, which is why it keeps working when everything above it has\nfailed. The intelligence layer's triggers are evaluated against bus events rather than\nagainst protocol internals.\n\n## Data flow\n\n```\nSentinel ──publish(metrics, medium)──┐\n         └─publish(anomaly, high)────┤\n                                     ▼\n                              AsyncEventBus\n                                     │  (priority ordered, bounded)\n                 ┌───────────────────┼───────────────────┐\n                 ▼                   ▼                   ▼\n            Serpentine        intelligence triggers   audit record\n```\n\n## Tradeoffs\n\n**Indirection costs traceability.** A direct call stack shows you what invoked what. A\nbus does not, and reconstructing a causal chain afterwards requires the events to carry\nenough context to do it. This is part of why decisions carry their own provenance —\nsee [explainability before automation](/vector/explainability-before-automation).\n\n**Priority can starve.** A sustained flood of high-priority events delays medium-priority\nones indefinitely. Acceptable here because a sustained anomaly flood is itself a\ncondition the stack is designed to escalate on, but it is a real property rather than\nan oversight.\n\n**Bounded queues drop.** At the limit, something is refused. Refusal at a known\nboundary is preferable to unbounded growth, and it is still a loss.\n\n## Failure modes\n\n**Queue saturation.** Backpressure at `max_queue_size`. Bounded and observable.\n\n**Subscriber slower than publisher.** The queue grows toward its bound; the slow\nconsumer is the fault, and the bus makes it visible rather than absorbing it silently.\n\n**A handler raising.** Whether one failing subscriber affects delivery to others is not\nevident from the structure alone, and this document does not assert an answer.\n\n## Future evolution\n\nThe natural direction is per-topic backpressure policy — allowing telemetry to be shed\nunder load while alarms are never dropped. That is inference about a sensible next step,\nnot a documented plan.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-event-bus"
    },
    {
      "kind": "document",
      "slug": "the-hardware-abstraction-layer",
      "title": "The hardware abstraction layer",
      "summary": "One interface between the decision system and a physical signal controller — and why emergency override lives at this layer rather than above it.",
      "body": "*[VECTOR archive](/vector) › Architecture › Hardware abstraction*\n\nEverything above this layer reasons about traffic. Below it, something changes a light.\nThe boundary between those two is the point where a decision stops being a\nrepresentation and becomes an action in the world, and it is drawn deliberately narrow.\n\n## Problem\n\nA decision system coupled to a specific controller cannot be tested without that\ncontroller, cannot be deployed onto different hardware, and cannot be run in simulation.\nThe coupling also spreads: once one component knows the device, others start assuming\nits behaviour.\n\n## Context\n\nThe system must run against simulated controllers during development, against a\nsimulation environment for training, and eventually against physical infrastructure —\nwithout the decision path knowing which.\n\n## Constraints\n\nThe interface has to be small enough that implementing it for new hardware is\ntractable, and expressive enough to carry an emergency stop. It must be synchronous at\nthe point of action: a signal change is not something to be queued.\n\n## Alternatives considered\n\n**Direct control from the decision engine.** Fastest path, no indirection, and it makes\nthe decision engine untestable without hardware.\n\n**A general device abstraction.** More flexible, considerably larger, and generality\nthat no second device type has yet demanded is generality bought on speculation.\n\n## Chosen architecture\n\n`SignalController` in `hardware/controller_interface.py` — an abstract base class with\nthree operations:\n\n- `set_phase(phase, duration)` — the entire ordinary control surface\n- `get_state()` — what the controller currently reports\n- `emergency_override()` — a distinct operation, not a special case of `set_phase`\n\nThe narrowness is the design. Three methods is a small enough contract that\n`simulated_controller.py` is a faithful stand-in rather than an approximation, which is\nwhat makes testing against it meaningful.\n\n**Emergency override is a separate method**, and that is the most consequential\ndecision here. Expressing an emergency as a phase change with particular arguments would\nroute it through the same validation, queuing and interpretation as an ordinary\ncommand — the path most likely to be congested or degraded when it is needed. A distinct\noperation can take a distinct path. This is the device-level counterpart of\n[Hyperion](/vector/protocol-hyperion): the emergency mechanism is deliberately not the\nordinary mechanism with an urgent flag.\n\n`controller_worker.py` and `control_loop.py` sit around this interface, isolating the\ntiming of physical control from the decision cycle above it.\n\n## Subsystem relationships\n\nConsumes decisions from the decision engine in `core/`. Implemented by\n`simulated_controller.py` for development, and paired with\n[the simulation layer](/vector/the-simulation-layer) for training. Its `get_state()`\noutput is part of what [Sentinel](/vector/protocol-sentinel) ultimately observes.\n\n## Data flow\n\n```\ndecision → control_loop → controller_worker → SignalController.set_phase()\n                                            ↘ emergency_override()  (separate path)\ndevice state ← SignalController.get_state() ← ─────────────────────────────────\n```\n\n## Tradeoffs\n\n**A narrow interface constrains what can be expressed.** Anything a controller can do\nbeyond phase and duration is not reachable. Accepted deliberately: a wider interface is\nharder to implement faithfully, and an unfaithful simulated implementation makes every\ntest that uses it misleading.\n\n**Abstraction hides device-specific failure.** A controller that fails in a way the\ninterface has no vocabulary for surfaces as a generic error, if at all.\n\n## Failure modes\n\n**The controller does not respond.** The decision was made and the world did not change.\nWhether the layer distinguishes \"command accepted\" from \"command took effect\" is not\nevident from the interface, and it is the distinction that matters most here.\n\n**State drift.** The reported state and the physical state diverge, and every layer\nabove is reasoning about a controller that does not exist as described.\n\n**Override unavailable.** The emergency path is the one that must work when the ordinary\npath does not; that it is separate is the mitigation, and it is not a guarantee.\n\n## Future evolution\n\nCommand acknowledgement — confirming effect rather than acceptance — is the most\nvaluable addition, and it would give the layers above a real signal for the difference\nbetween a decision made and a decision applied.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-hardware-abstraction-layer"
    },
    {
      "kind": "document",
      "slug": "the-human-remains-the-operator",
      "title": "The human remains the operator",
      "summary": "The final principle, and the one the others exist to make possible. Authority is delegated, bounded, and returns to a person by default.",
      "body": "Autonomy is usually described as a property a system has. It is more accurate to\ndescribe it as authority a person has delegated — because that framing makes the two\nimportant questions unavoidable: how much, and how is it taken back.\n\n**The principle: the operator is a person. The system exercises delegated authority\nwithin a boundary that person defined, and cannot widen it.**\n\n## What the boundary is\n\nThe constraints under which the system operates are written by a human, in a form a\nhuman can read and argue with, and the system cannot modify them. This is not a\nlimitation that was worked around; it is the definition of the arrangement. A system\nable to revise its own constraints does not have constraints — it has preferences.\n\nThe corollary is that the constraint layer is deliberately not learned. Everything else\nin the system may be adaptive. The boundary may not, because a boundary that moves in\nresponse to the pressure applied to it is not a boundary.\n\n## Autonomy is earned, not assumed\n\nThe delegation is conditional and continuously re-evaluated. Under sustained\nuncertainty the system narrows its own authority — disabling its more speculative\ncomponents, falling back to the deterministic floor — and restores that authority only\nin stages, against evidence of stability rather than the passage of time.\n\nThis inverts the intuitive response. A system in difficulty appears to need more\ncapability brought to bear. In practice, difficulty means conditions have left the\nrange the models were built for, so the moment intelligence is most wanted is the\nmoment it is least reliable. Narrowing under uncertainty is the correct direction, and\nit is uncomfortable to implement precisely because it feels backwards.\n\n## The end of the chain must be unreachable from inside\n\nThe escalation chain ends at a permanent termination that the system cannot invoke.\nOrchestrators, internal services and automated components are explicitly blocked from\nreaching it; it requires a person, physically present, with multi-factor authorisation,\nand it preserves a full state archive before ending anything.\n\nThat the system cannot reach the end of its own escalation chain is the single most\nimportant structural property in the architecture. Everything above that point is the\nsystem managing itself. The final step is where it is not permitted to.\n\nBuilding that capability at all required treating the ability to stop as a feature to\nbe designed rather than an absence to be tolerated —\n[I stopped optimising, and built a way to stop](/vector/we-built-a-way-to-stop).\n\n## Why this principle is last\n\nBecause the others exist to make it possible.\n\nA person can only overrule a system whose decisions they can understand, which requires\n[explainability](/vector/explainability-before-automation). Explanation requires\nobservations that arrived with their provenance intact, which requires\n[observation before action](/vector/observation-before-action). Meaningful intervention\nrequires a defined state to intervene toward, which requires\n[governance before intelligence](/vector/governance-before-intelligence). And the\nreasoning behind the boundary has to be written down, or the next person to hold it\ninherits a rule with no argument attached — which is how boundaries erode.\n\nA system that satisfies the first five principles and not this one is well-engineered\nand unowned. The purpose of the Canon is that it never becomes that.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-human-remains-the-operator"
    },
    {
      "kind": "document",
      "slug": "the-intelligence-layer",
      "title": "The intelligence layer",
      "summary": "Anarchy, ORACLE, MORL, PULSAR, Fusion and Meta — five ways of being clever and one controller that decides how much cleverness is currently affordable.",
      "body": "With the governance stack proven, the intelligence layer was built on top of it. The\nfull pipeline:\n\n```\nInput → Sentinel → Serpentine → Anarchy → ORACLE → MORL → PULSAR → Fusion → Meta → Output\n```\n\nNote where it starts. Detection and classification run *before* any intelligence does.\nThe system knows how stressed it is before it decides how ambitious to be.\n\n## Anarchy — bounded simulation\n\nEvent-triggered rather than continuous. Anarchy wakes on specific triggers — traffic\nspike, error surge, confidence collapse, novelty detection, strategy failure loop —\nand spawns a lightweight micro-simulator to evaluate three candidate strategies:\nbaseline (what the system was going to do), conservative (stability-focused) and\naggressive (performance-focused).\n\nEach is scored:\n\n```\nScore = (Accuracy × 0.4)\n      + (Stability × 0.3)\n      - (Latency × 0.15)\n      - (Resource Cost × 0.15)\n```\n\nAn override fires only if the best simulated score beats baseline by at least 10%, and\nthe whole thing runs inside a hard 30 ms budget.\n\nBoth numbers are doing real work. The 10% margin means marginal improvements do not\nget to disturb a working system — the default has to be beaten decisively, not\nnarrowly. The 30 ms cap means the deliberation cannot become the emergency.\n\n## ORACLE — the world model\n\nAn RSSM-based world model that predicts future states in latent space rather than\nrunning full simulations. Short-path conservative mode, confidence scaling, and\nmemory-assisted candidate ordering.\n\nThe system imagines futures instead of simulating them, which is the only version that\nfits inside the time budget. A simulator that is accurate and too slow produces\nnothing; the prediction has to arrive while the decision is still open.\n\n## MORL — competing objectives\n\nMulti-objective reinforcement learning across goals that genuinely conflict: minimise\ntravel time, minimise emissions, prioritise emergency vehicles, reduce fuel\nconsumption. Every one of those trades against at least one other.\n\nMORL is automatically gated off under low confidence or high risk. The most\nsophisticated component is the first one switched off when conditions degrade, which\nis the correct ordering and an uncomfortable one to implement.\n\n## PULSAR — small, reversible corrections\n\nA deterministic micro-adjustment engine. Reversible tuning under hard pulse-count and\nlatency budgets. Small, precise, bounded.\n\nPULSAR exists because not every correction should go through the full decision path.\nSome adjustments are small enough that the deliberation costs more than the error, and\nreversibility is what makes it safe to skip the deliberation.\n\n## Fusion — the coordinator\n\nCoordinates ORACLE, MORL and PULSAR under safety gates: trust recalibration, emergency\noracle down-weighting, a safety fallback gate, and authority-aware clamps.\n\nFusion is where \"how much do I currently believe each of my own components\" is decided,\nand it is answered continuously rather than configured once.\n\n## Meta — slow adaptation\n\nA background controller that monitors runtime signals and applies smoothed control\nupdates — but only when stable conditions justify adapting at all. Gradual,\nconservative, never aggressive.\n\nMeta is the layer that could most easily destroy the system, because a controller that\nadapts the controller during instability amplifies exactly what it is trying to damp.\nSo it is deliberately the slowest thing in the architecture, and it refuses to act\nwhile the system is stressed.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-intelligence-layer"
    },
    {
      "kind": "document",
      "slug": "the-latency-crisis",
      "title": "The latency crisis: one component, ninety-seven percent",
      "summary": "p95 of 100 ms to 1.124 ms. Six weeks of looking in the wrong place, and the component I trusted most turning out to be the entire problem.",
      "body": "After the intelligence layer was built, the system produced brutal latency numbers. p95\naround 100 ms. Tail spikes. Instability under stress. Not remotely real-time capable.\n\nThis is the most useful thing that has happened to the project, so it is written down\nin more detail than it strictly needs.\n\n## Looking in the wrong place\n\nThe obvious suspects were the new components. ORACLE runs a world model. MORL does\nmulti-objective optimisation. Anarchy spawns simulations. Those are the expensive-\nsounding things, and I spent weeks assuming the cost was where the sophistication was.\n\nIt was not. I profiled everything properly, and Sentinel — the anomaly detector, the\nsimplest component, the one written first and trusted longest — was consuming roughly\n**97% of runtime**.\n\nIt was the component I had never suspected precisely because it was the one I\nunderstood best.\n\n## Why it was so expensive\n\nSentinel ran heavy statistical checks on every cycle, with unbounded work patterns.\nUnbounded is the operative word: the cost varied with conditions, so it was not a\nconstant overhead that shows up cleanly in an average. It amplified into tail latency\nspikes, and because Sentinel sits at the head of the pipeline, every spike poisoned\neverything downstream.\n\nEvery stage after it inherited the delay and reported its own latency honestly, which\nmade the whole pipeline look uniformly slow rather than pointing anywhere.\n\n## The fix\n\n- Rolling window metrics with O(1) incremental statistics\n- A 5 ms hard cap on Sentinel's time budget\n- Heavy analysis moved to bounded background tasks\n- Strict per-stage budget enforcement across the entire pipeline\n- Early-exit logic when risk or latency indicators exceed thresholds\n- Spike detection with temporary mitigation windows afterwards\n\nOne bottleneck. One fix. The system went from unstable to real-time.\n\n## The numbers\n\n| Phase | p95 latency | Stability |\n| --- | --- | --- |\n| Early profile | ~100 ms | Brittle |\n| Post-optimisation | ~2 ms | Stable |\n| Current (500-cycle snapshot) | 1.124 ms | Stable |\n\nFull snapshot: 0.607 ms average, 1.124 ms p95, 1.634 ms p99, 1.908 ms maximum. Zero\ndropped events on the bus, zero ordering violations, 32 of 32 regression tests passing.\n\n## What I actually learned\n\n**Find the bottleneck before optimising anything.** Everything else was noise. Six\nweeks of careful work on components that were not the problem produced nothing, and\nwould have produced nothing however well I had done it.\n\n**Suspicion should follow measurement, not architecture.** I searched where the\ncomplexity was, because complexity feels expensive. Cost does not care how\nsophisticated a component looks.\n\n**Unbounded is the defect, not slow.** Sentinel was not slow on average. It was\nunbounded, which meant its worst case was unrelated to its typical case — and in a\nreal-time system the worst case is the specification. This is why every stage now\ncarries a hard budget.\n\n**Tail risk is a first-class problem.** Average latency means nothing if p99 is\ncatastrophic. The early system had a tolerable average and an unusable tail.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-latency-crisis"
    },
    {
      "kind": "document",
      "slug": "the-name",
      "title": "The name",
      "summary": "EXTREMIS to VECTOR. Why a name is a constraint on the thinking rather than a label on the result, and how to notice when yours has become the wrong one.",
      "body": "Naming is usually treated as the last and least consequential decision in a project.\nIt is neither. A name determines the vocabulary available for describing the thing,\nand vocabulary determines which designs occur to you — silently, without ever\nproducing an error that could be noticed.\n\nThis document records a rename, and the general lesson is more useful than the\nspecific history.\n\n## EXTREMIS\n\nThe project was called EXTREMIS from its conception in October 2025 until the middle of\n2026. The name came, roughly, from Marvel, and was chosen when the project was an\nadaptive traffic system and nothing more.\n\nIt was a good name for that. It was evocative, it suggested response under duress, and\nit fit a system whose job was to react to conditions at the edge.\n\n## What changed underneath it\n\nThe system stopped being an adaptive traffic controller. It became a fault-tolerant\nprotocol stack, then a real-time system with an intelligence layer inside it, then a\nbounded and governed structure whose central concern was the authority of the thing\nrather than the quality of its decisions.\n\nThe name did not move. It described the first version accurately and every subsequent\nversion less so, and because a name does not throw errors, nothing announced the\ngrowing mismatch.\n\n## How the mismatch made itself known\n\nNot as dissatisfaction with the name. As a pattern in explanations.\n\nEvery attempt to describe a general component began at traffic and reached outward\napologetically. Explaining the termination protocol would start with intersections,\neven though termination has nothing to do with intersections, because traffic was the\nvocabulary the name had established and general statements had to be translated out of\nit.\n\nThat is a communication symptom of an architectural problem. Components were being\ndesigned against traffic-shaped assumptions, because those were the assumptions the\navailable vocabulary made easy to hold. Generalising a component afterwards costs\nconsiderably more than never having specialised it.\n\nThere was also a trivial reason, which is that the name belongs thoroughly to Iron\nMan. On its own it would have been an annoyance rather than a migration.\n\n## VECTOR\n\nVariable Environment Control Through Observation Response.\n\nIt is an acronym and also a description. A vector carries a magnitude and a direction,\nand the interesting property of an autonomous decision is never how much it did but\nwhich way it pointed and why.\n\nWhat it does structurally is name no domain. A component designed under this name has\nto state its assumptions rather than inherit them, because the name supplies none.\nThat is the entire benefit, and it is worth the cost of a full migration.\n\n## The general rule\n\n**A name is a constraint on the thinking, not a label on the result.**\n\nThe signal that a name has become wrong is not boredom with it. It is noticing that\nexplanations of general things keep starting from a specific case. When that pattern\nappears, the vocabulary has narrowed the design space, and the cost of continuing is\npaid in components that have to be rewritten later.\n\nRename early. The migration is mechanical; the accumulated specialisation is not.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-name"
    },
    {
      "kind": "document",
      "slug": "the-philosophy",
      "title": "The philosophy",
      "summary": "The six commitments encoded in the name — observation, response, variable environments, cognition, documentation, provenance — and why each one is there.",
      "body": "A name that is an acronym usually means nothing beyond its expansion. This one was\nchosen after the fact to describe commitments the system had already made, and each\nterm answers a question that had to be settled before anything could be built.\n\n## Why observation\n\nBecause the alternative is assumption. A system that acts on a schedule has assumed a\nworld; a system that acts on observation has looked at one.\n\nThe commitment is stronger than \"read sensors\". An observation carries how reliable it\nis, alongside what it says. A measurement stripped of its uncertainty is a number\npretending to be a fact, and every downstream safety property depends on knowing the\ndifference. This is developed in\n[observation before action](/vector/observation-before-action).\n\n## Why response\n\nBecause \"control\" would have been the wrong word. Control implies the system\ndetermines the state of the environment. It does not. It responds to an environment\nthat has its own causes, most of which the system cannot see and none of which it\ncommands.\n\nThe distinction matters in failure. A system that believes it controls its environment\ninterprets an unexpected state as an error to be corrected harder. A system that knows\nit is responding interprets the same state as evidence that its model is wrong — and\nthe second interpretation is the one that does not escalate.\n\n## Why variable environments\n\nBecause stationarity is the assumption that quietly invalidates everything else. Most\nengineering methods assume the distribution generating your data is the distribution\nyou will meet. Real environments — traffic, weather, human behaviour — are not\nstationary. They drift, they have regimes, and they occasionally do something with no\nprecedent.\n\nBuilding for a variable environment means the system must treat \"conditions have left\nthe range I understand\" as an expected event with a defined response, rather than as\nan anomaly to be handled later.\n\n## Why cognition\n\nBecause the system must hold a representation of the environment rather than merely\nreact to it. Reaction is stateless: an input arrives, an output is produced, nothing\nis retained. That is sufficient for a thermostat and insufficient here, because the\nright action frequently depends on what has been happening rather than on what is\nhappening.\n\nCognition in this archive is a precise and modest claim. The system maintains a model\nof its environment, predicts forward from it, and evaluates its own predictions\nagainst what occurred. It is not a claim about understanding, and nothing in this\narchive should be read as one.\n\n## Why documentation\n\nBecause a decision whose reasoning was never recorded is re-litigated every six months\nby someone with less context, usually the person who made it.\n\nThe stronger version, and the one actually practised: writing the argument for a\ncomponent before building it is the cheapest way to discover that it should not exist.\nDocumentation is not a description produced after the engineering. It is where a\nmeaningful fraction of the engineering happens. See\n[documentation is part of engineering](/vector/documentation-is-part-of-engineering).\n\n## Why provenance\n\nBecause \"where did this come from\" is a question that gets asked exactly when it can\nno longer be answered.\n\nEvery value moving through the system has an origin, a confidence, and a path. A\ndecision that cannot say which observations it rested on cannot be audited after an\nincident, and a model whose training provenance is unrecorded cannot be reasoned about\nwhen it starts behaving oddly. Provenance is expensive, unglamorous, and the thing\nthat makes an incident investigable rather than merely regrettable.\n\nThe same principle governs this archive. Documents record whether they were\nAI-assisted and whether a human reviewed them, and publication is refused when the\nanswer is unsatisfactory — because a claim about provenance that nothing enforces is\na claim, not a property. How that is enforced in the software you are reading is\ndescribed in [provenance belongs in the schema](/vector/provenance-as-a-schema).",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-philosophy"
    },
    {
      "kind": "document",
      "slug": "the-protocol-stack",
      "title": "The eight-protocol governance stack",
      "summary": "Sentinel to Terminus: one escalation chain from the first anomaly to a permanent, human-authorised end. Every protocol has one role and nothing overlaps.",
      "body": "The governance stack is eight protocols forming a single escalation chain. It was the\nfirst major milestone of the project and it was built before any intelligence sat on\ntop of it, because a system that cannot govern itself has no business being clever.\n\n```\nSentinel     →  detects anomalies\nSerpentine   →  classifies stress state\nAegis        →  contains failures\nAtlas        →  rebalances load\nMender     →  repairs with AI\nPhoenix      →  orchestrates recovery\nHyperion     →  autonomous emergency shutdown\nTerminus     →  permanent human-authorised termination\n```\n\nEvery protocol has a defined role. Every handoff is deliberate. Nothing overlaps — and\nthe non-overlap is a design rule rather than an outcome, because overlapping\nresponsibilities are where escalation chains rot.\n\n## Sentinel — detection\n\nRuns every second, collecting latency, CPU, memory, GPU, queue depth and throughput.\nIt uses z-score analysis against a rolling history of 20 samples to catch anomalies\nbefore they become failures. Every anomaly is categorised and severity-scored.\n\nRaw metrics travel at medium priority. Detected anomalies escalate at high priority\nimmediately — the bus does not treat \"something is wrong\" as ordinary traffic.\n\nSentinel is also the component that nearly killed the project. See\n[the latency crisis](/vector/the-latency-crisis).\n\n## Serpentine — classification\n\nSynthesises everything Sentinel produces into one unified stress state across five\nlevels: stable, moderate, critical, failure, catastrophic. It derives them from three\nindices — pressure, congestion and instability.\n\nCatastrophic requires 92% pressure or 95% instability. Not lower. The thresholds are\ndeliberately unkind because a classification that fires early is a classification\nnobody believes, and an escalation chain nobody believes is decoration.\n\n## Aegis — containment\n\nStops failures cascading. Nodes above a 20% error rate are isolated. Nodes above 95%\nload are throttled. Global circuit breakers activate under catastrophic state. Every\naction is logged with explicit reasoning attached.\n\n## Atlas — rebalancing\n\nRedistributes load. Overloaded nodes above 80% shift work to underutilised nodes below\n50%, targeting 65% after the shift. When redistribution is not enough, Atlas\nrecommends autoscaler activation rather than pretending it solved the problem.\n\n## Mender — repair\n\nA dual-head neural network diagnoses incidents and ranks repair strategies by risk.\n\nThe interesting part is the governor rather than the network: an adaptive credit\nsystem, regenerating 0.02 credits per second to a maximum of 10, which prevents\nthrashing. As failures accumulate, credits deplete and the system becomes\nprogressively more conservative. It cannot enter a loop of aggressive repairs, because\naggression is a budget and the budget runs out.\n\n## Phoenix — recovery\n\nBrings the system back after repair. Sequenced restarts in dependency order. Container\nrebuilds for nodes above 85% load. Global runtime state restoration.\n\nConservative by design: it performs more restarts than strictly necessary, never\nfewer. Restarting something that did not need it costs milliseconds. Not restarting\nsomething that did costs the recovery.\n\n## Hyperion — emergency shutdown\n\nFires within one second. Halts everything. No complex logic, no second-guessing,\nirreversible.\n\nHyperion is deliberately the least sophisticated component in the system. Anything\nclever here is another thing that can be wrong at the exact moment nothing else is\nworking.\n\n## Terminus — the human kill switch\n\nTerminus blocks orchestrators, AI agents and internal services from invoking it. It\nrequires a human at a physical console with multi-factor authorisation. It encrypts a\nfull system state archive before terminating.\n\nRecovery afterwards is possible — but only deliberately, and only by a person.\n\nThat the system cannot reach the end of its own escalation chain is the most important\nproperty in this document. Everything above Terminus is the machine managing itself.\nTerminus is the point where it is not allowed to.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-protocol-stack"
    },
    {
      "kind": "document",
      "slug": "the-simulation-layer",
      "title": "The simulation layer",
      "summary": "Why a system intended for physical infrastructure is developed against a simulator, and what a simulator can and cannot establish.",
      "body": "*[VECTOR archive](/vector) › Architecture › Simulation*\n\nA control policy for traffic infrastructure cannot be learned on traffic infrastructure.\nThe exploration required to learn anything is precisely what must never happen on a\nlive road, which makes simulation a requirement rather than a convenience.\n\n## Problem\n\nLearned components need many more interactions than reality can safely provide, and the\ninteractions they need most are the unusual ones. A system that only ever observes\nnormal operation has no basis for behaving well outside it.\n\n## Context\n\nSUMO — Simulation of Urban Mobility — is an established open-source traffic simulator.\nUsing it rather than building a simulator means inheriting a model of traffic that was\nnot written to make this system look good, which is a real epistemic advantage.\n\n## Constraints\n\nThe interface the simulator presents has to be the same shape the system sees in\nproduction, or the policy learns against an environment it will never meet. Rollouts\nmust be recordable for training. And the dependency has to be replaceable — binding the\nsystem to one simulator is the coupling problem again, one layer out.\n\n## Alternatives considered\n\n**A purpose-built simulator.** Full control, and it would encode the same assumptions as\nthe system being tested. A simulator written by the author of the controller tends to\nreward the controller.\n\n**Recorded traffic data replayed.** Real, and non-interactive: it cannot answer what\nwould have happened had the system acted differently, which is the entire question.\n\n## Chosen architecture\n\n`BaseSimulationAdapter` in `simulation/sumo_adapter.py` defines a familiar reinforcement\nlearning environment contract — `reset`, `get_state`, `get_valid_actions`, `step` — with\n`SumoAdapterConfig` carrying configuration and the SUMO binding behind it.\n\n`get_valid_actions` is worth noting. The environment states which actions are\npermissible rather than leaving the agent to propose something invalid and be corrected,\nwhich is the same relationship the policy layer has to the decision path in the live\nsystem: constraints available during the decision rather than as a veto afterwards. See\n[governance before intelligence](/vector/governance-before-intelligence).\n\nRollouts are captured as `TransitionRecord` and exposed as a `RolloutDataset`, so\nsimulated experience enters training through a typed boundary rather than an ad-hoc one.\n\n## Subsystem relationships\n\nProduces training data for the learned components in `learning/`, entering training\nthrough the same typed boundary that historical data does — see\n[the data pipeline](/vector/the-data-pipeline). Mirrors the contract of\n[the hardware abstraction layer](/vector/the-hardware-abstraction-layer), which is what\nallows the same decision path to run against a simulator or a controller.\n\n## Data flow\n\n```\nreset → get_state → get_valid_actions → (policy) → step → TransitionRecord\n   ↑                                                            │\n   └──────────────────── episode ends ──────────── RolloutDataset → training\n```\n\n## Tradeoffs\n\n**A simulator is a model, and every model is wrong somewhere.** Policies learned in\nsimulation encode its inaccuracies. This is the central limitation and no amount of\nsimulation fidelity removes it.\n\n**Simulated experience is cheap, which makes it tempting to trust.** Large volumes of\nsimulated data produce confident estimates about a world that is not the one the system\nwill operate in.\n\n## Failure modes\n\n**Reality gap.** Behaviour that is correct in simulation and wrong in deployment. The\nruntime defence is not better simulation — it is that\n[the authority layer](/vector/the-authority-layer) narrows autonomy when conditions stop matching\nexpectations.\n\n**Overfitting to the scenario set.** A policy excellent on the scenarios it was trained\nagainst and undistinguished elsewhere.\n\n## Future evolution\n\nA pilot against real infrastructure is the outstanding item, and it is the only thing\nthat can establish what simulation cannot — noted as unfinished in\n[the v1 changelog](/vector/changelog-v1). A digital twin, running alongside the live\nsystem rather than ahead of it, is the longer-term direction.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/the-simulation-layer"
    },
    {
      "kind": "document",
      "slug": "vision",
      "title": "Vision",
      "summary": "Where this is going over five, ten and twenty years — as engineering direction rather than ambition, and with the uncertainty left in.",
      "body": "A vision document is usually a description of a desired outcome. This one is a\ndescription of a direction, which is a weaker and more useful thing: it states what\nthe work is converging on without claiming to know where it arrives.\n\nEverything here should be read as a current best estimate that later documents are\npermitted to contradict. When one does, the contradiction stays visible.\n\n## Five years — one domain, properly\n\nThe nearest goal is unglamorous: VECTOR operating real traffic infrastructure, in one\nplace, continuously, with its failures documented rather than smoothed over.\n\nThat requires things the system does not yet have. Numbers produced under simulation\nand adversarial testing describe a laboratory; until a system has met an environment\nthat did not read its assumptions, its measurements are provisional. It requires a\ndefence against sensor inputs that are wrong on purpose rather than by accident,\nwhich is a category the current design acknowledges and does not fully address. And it\nrequires the governance layer to be exercised by real incidents, because an escalation\nchain that has never escalated under genuine pressure is a hypothesis.\n\nThe measure of success at five years is not performance. It is whether the accountability\nstructure held when something went wrong, and whether the record it produced was\nsufficient to explain the event to someone who was not there.\n\n## Ten years — the domain stops mattering\n\nThe architecture is arranged so that the domain is a component rather than an\nassumption. If that arrangement is correct, VECTOR should apply to any environment\nwith the same shape: continuously changing, partially observable, consequential to get\nwrong, and currently run either by a fixed schedule or by a person who cannot watch it\nconstantly.\n\nEnergy distribution, water systems, and industrial process control all have that shape.\nWhether the abstraction genuinely holds is an open question — this is exactly the kind\nof claim that is easy to make and expensive to test, and the honest position is that a\nsingle second domain would tell me more than a decade of reasoning about it.\n\nThe failure mode to watch for is a generality that exists only in the documentation.\nIf moving to a second domain requires changing the governance layer, the layer was\nnever general, and this archive should say so plainly when it happens.\n\n## Twenty years — the part that is not software\n\nThe long-term question is not technical. Systems that act on the physical world with\npartial autonomy are going to become ordinary, and the practices for holding them\naccountable are currently improvised per project.\n\nWhat would be worth building over that horizon is a set of conventions rather than a\nproduct: what an autonomous system owes its operator, what a decision record has to\ncontain to be worth keeping, what it means for a termination path to be genuinely\nreachable, and what evidence a system should have to produce before being trusted with\na wider envelope.\n\nVECTOR is one attempt at answering those questions in a specific case. The attempt is\nmore likely to be useful than the artefact.\n\n## What this vision deliberately excludes\n\nNo claim about scale, adoption or organisational form appears here, because none of\nthose are engineering direction and including them would date this document within a\nyear.\n\nNo timeline appears beyond the three horizons above, and those are approximate. A\ndated roadmap belongs in the changelog, where being superseded is the expected outcome\nrather than an embarrassment.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/vision"
    },
    {
      "kind": "document",
      "slug": "we-built-a-way-to-stop",
      "title": "I stopped optimising, and built a way to stop",
      "summary": "Performance describes a system under favourable conditions. Infrastructure only matters under the other ones.",
      "body": "For a long time progress looked obvious. Models got better. Outputs stabilised. Loss\nwent down. The system appeared to work.\n\nThat kind of progress is deceptive. Performance tells you how a system behaves when\nconditions are favourable. It says almost nothing about how it behaves when it is\nwrong, uncertain, or under stress — and in real infrastructure, those are the only\nmoments that matter.\n\nSo at one point I stopped optimising. Not because the system had failed, but because\nsuccess without restraint is just delayed failure.\n\n## The uncomfortable realisation\n\nMost intelligent systems are designed around a single obsession: do more, faster,\nbetter. That works until the environment stops cooperating, and then a system with no\nconcept of restraint simply continues — confidently, at speed, in the wrong direction.\n\nThe capability that was missing was not a better decision. It was the ability to stop\nmaking them.\n\n## What stopping requires\n\nIt sounds like the easy feature and it is the hard one, because stopping well means\nanswering questions that optimisation never asks.\n\nWhat state is the system left in? Who is allowed to invoke it, and who is explicitly\nnot? Is it reversible, and if so by whom? What is preserved for the people who have to\nwork out afterwards what happened?\n\nThose questions produced the top of the escalation chain. Hyperion halts everything\nwithin a second with no cleverness at all. Terminus is unreachable by orchestrators,\nagents and internal services — it requires a human at a physical console with\nmulti-factor authorisation, and it encrypts a full state archive before it ends\nanything.\n\n## The principle\n\n**A system that cannot stop is not autonomous. It is unsupervised.**\n\nThe difference is whether stopping is a designed capability with an owner, or an\nabsence nobody planned for. Everything in VECTOR that looks like caution — the credit\nsystem in Mender, the 10% override margin in Anarchy, the staged recovery in the authority layer —\ndescends from taking that seriously.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/we-built-a-way-to-stop"
    },
    {
      "kind": "document",
      "slug": "what-is-vector",
      "title": "What is VECTOR?",
      "summary": "The preface. What the system is as an idea, before any question of how it is implemented — and what it is deliberately not.",
      "body": "Every engineered system embodies assumptions about the world in which it operates.\nTraditional traffic infrastructure assumes that yesterday's timing schedule is a\nreasonable approximation of today's conditions. VECTOR begins by rejecting that\nassumption, and then spends most of its structure dealing with the consequences of\nhaving rejected it.\n\nVECTOR — Variable Environment Control Through Observation Response — is a system for\noperating infrastructure whose conditions change faster than a person can respond to\nthem, without surrendering the accountability that a person operating it would\nprovide.\n\n## The idea, stated once\n\nA system that observes an environment and acts on it has been granted authority. That\nis true whether or not anyone decided to grant it. The moment a mechanism can change\nthe state of the world in response to what it perceives, someone is responsible for\nwhat it does, and the question that matters is whether the system is built so that\nresponsibility can actually be exercised.\n\nMost of the difficulty in autonomous infrastructure is not perception and not\ndecision-making. It is that these three requirements pull against each other:\n\n- The system must act well, which pushes toward sophistication.\n- The system must be able to account for what it did, which pushes against\n  sophistication, because the most capable methods are frequently the least legible.\n- A person must be able to overrule it, which requires the first two to hold — you\n  cannot meaningfully overrule a decision you cannot understand.\n\nVECTOR is an attempt to satisfy all three at once. Everything else in this archive is\na consequence of that attempt.\n\n## What it is not\n\n**It is not a model.** A model maps inputs to outputs under conditions resembling\nthose it was trained on. VECTOR contains models. It is not one, and the distinction\nis the subject of [why VECTOR exists](/vector/why-vector-exists).\n\n**It is not autonomous in the ordinary sense.** Its authority has a boundary, that\nboundary is written by a person, and the system cannot revise it. A system able to\nrewrite its own constraints does not have constraints.\n\n**It is not a traffic product.** Traffic is the first application and remains the one\nthe system is tested against. It is not the subject. The subject is the general\nproblem above, and traffic is where that problem is concrete enough to be worked on.\n\n**It is not finished.** This archive documents a system under construction, including\nthe parts that were wrong for months at a time. Where something is uncertain, the\ndocument says so.\n\n## How to read this archive\n\nThe five documents of Part I establish what VECTOR is and why: this preface,\n[why VECTOR exists](/vector/why-vector-exists),\n[the philosophy](/vector/the-philosophy), [the name](/vector/the-name), and\n[the vision](/vector/vision). The six that follow them state the principles every\ncomponent is required to obey. Together they are the Canon, and every later document —\narchitecture, protocol specification, research note, failure report — assumes they have\nbeen read and does not repeat them.\n\nRead in order, the archive is an argument rather than a reference manual. It can be\nused as a reference; it was not written as one.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/what-is-vector"
    },
    {
      "kind": "document",
      "slug": "why-vector-exists",
      "title": "Why VECTOR exists",
      "summary": "Static infrastructure fails in one way, adaptive infrastructure fails in another, and intelligence alone resolves neither. What is actually missing is governance.",
      "body": "A system can fail by being wrong. It can also fail by being right about a world that\nhas stopped existing. Infrastructure fails almost entirely in the second way, and that\ndistinction is why [VECTOR](/vector/what-is-vector) is built the way it is.\n\n## Why static infrastructure fails\n\nA fixed-cycle traffic signal is not a bad piece of engineering. It is predictable,\nauditable, cheap to reason about, and extremely difficult to make worse — genuine\nvirtues when the failure mode involves vehicles. Any argument for replacing it has to\nbegin by taking those virtues seriously, and most do not.\n\nIts failure is structural rather than incidental. The timing was derived once, from\nconditions observed at some past moment, and is then applied indefinitely. It does not\ndegrade when it becomes wrong, because it has no way of knowing that it has. A signal\nholding a red against an empty road at three in the morning is executing its\nspecification perfectly.\n\nThe cost is invisible and continuous. Nobody files a report about the ninety seconds\nthey waited unnecessarily, so the failure never accumulates into evidence, and the\nsystem is never revised because nothing ever appears to be wrong with it.\n\n## Why adaptive infrastructure fails differently\n\nThe obvious correction — make it respond to conditions — introduces a failure mode\nthat is worse in a specific way. A static system is wrong predictably. An adaptive\nsystem is wrong occasionally and unpredictably, and the occasions correlate with\nunusual conditions, which is precisely when the consequences are largest.\n\nA signal that has learned to optimise throughput will encounter a situation outside\nanything it has seen, and it will still produce an output, confidently, because\nproducing outputs is what it does. Nothing in the mechanism distinguishes \"this is a\nfamiliar situation and my answer is reliable\" from \"this is unlike anything I know and\nmy answer is a guess\".\n\nStatic systems fail loudly in aggregate and quietly per incident. Adaptive systems do\nthe reverse. Neither property is acceptable on its own, and the conditions under which\nadaptation is nevertheless the right choice are set out in\n[adaptive systems over static systems](/vector/adaptive-systems-over-static-systems).\n\n## Why intelligence alone is not the answer\n\nThe instinctive response is that the model needs to be better. Given sufficient data\nand capacity, the argument goes, the unusual situations stop being unusual.\n\nThis is true and insufficient, for a reason that does not depend on how good the model\ngets. Improving a model reduces how often it is wrong. It does nothing about what\nhappens when it is — and in infrastructure, the behaviour of the system during its own\nfailure is not a corner case. It is the specification.\n\nA system that is correct 99.99% of the time and undefined for the remainder has not\nbeen engineered. It has been tuned. The interesting question was never \"how often is\nit right\" but \"what does it do when it is not, and who finds out\".\n\nThis is also why the distinction between a model and a system is load-bearing. A model\nanswers: given this input, what is the correct output? A system answers a harder\nquestion: given that some inputs are stale, one sensor is lying, load is high, the\nlast three decisions did not have their predicted effect, and there are thirty\nmilliseconds — what happens now? No improvement to the first answer produces the\nsecond.\n\n## Why governance is the missing piece\n\nWhat is missing from both the static and the adaptive system is not capability. It is\nstructure around the capability: a defined answer for what the system does when it is\nuncertain, a boundary on what it is permitted to do at all, a record of why it did\nwhat it did, and a person with the standing to overrule it.\n\nThat structure is what this archive means by governance, and it is the actual subject\nof VECTOR. The intelligence is a component inside it. The governance is the product.\n\n## Why accountability is not optional\n\nA decision nobody can account for cannot be corrected, because correcting it requires\nknowing why it was made. It cannot be audited after an incident. It cannot be defended\nwhen it was right, which matters as much as being challenged when it was wrong.\n\nAccountability is therefore not an ethical addition to a working system. It is the\nproperty that makes the system improvable at all. A component that acts without leaving\nan account of its reasoning is a component that can never be debugged — only replaced.",
      "status": "PUBLISHED",
      "occurred": null,
      "recorded": null,
      "assisted": true,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/vector/why-vector-exists"
    },
    {
      "kind": "project",
      "slug": "attack-of-excalius",
      "title": "Attack of Excalius",
      "summary": "A Scratch platformer with hand-drawn sprites and a boss fight I never balanced. The first thing I finished.",
      "body": "I opened Scratch on 14 September 2020, at eight years old, and Attack of Excalius is\nwhat came out of it. A platformer. Every sprite drawn by hand, badly, with total\nconviction.\n\nIt is the oldest thing in this archive and it stays here permanently.\n\n## The problem\n\nI wanted to make a game. There was no problem statement, no user, and no requirement\nbeyond that — which, at eight, is exactly the right amount of specification.\n\n## The approach\n\nSprites drawn in the built-in editor. Jump physics arrived at by changing numbers\nuntil the jump felt right, which is a legitimate method I have never fully abandoned.\nCollision detection that worked in most cases and produced memorable failures in the\nrest.\n\nAnd a boss fight I could not balance. Every adjustment made it either trivial or\nimpossible, and I never found the middle. That was my first encounter with a problem\nthat does not yield to more effort of the same kind.\n\n## Architecture\n\nEvent-driven by construction, because Scratch gives you nothing else: sprites\nreceiving messages, each running its own scripts, all state global and shared.\n\nI did not know that \"every variable is global\" was a design constraint rather than how\nprograms are. I found out by breaking the game — a variable one sprite depended on,\nquietly modified by another, producing a bug I could not reason about because I had no\nconcept of scope to reason with. Meeting shared mutable state at eight, without the\nword for it, is a good way to learn why the word exists.\n\n## Results\n\nA finished game that other children played. Given how many projects at that age reach\n80% and stop, finishing is the whole result.\n\n## What I would do differently\n\n- **Finishing is a separate skill from building,** and it is the rarer one.\n- **The title screen is not the game.** I rewrote it more times than the boss fight,\n  because it was the part I knew how to improve. Working on what is comfortable instead\n  of what is hard is a habit I still catch myself in.\n- **Some problems do not yield to more of the same effort.** The boss fight needed a\n  different model of difficulty, not more tuning. I did not have that idea for years.",
      "status": "PUBLISHED",
      "occurred": "2020-09-14T00:00:00.000Z",
      "recorded": "2020-09-14T00:00:00.000Z",
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/projects/attack-of-excalius"
    },
    {
      "kind": "project",
      "slug": "mbot-neo",
      "title": "mBot Neo, and the end of robotics",
      "summary": "A robot that worked, and the discovery that I did not care. The most useful dead end I have hit.",
      "body": "Blockly and an mBot Neo, at eleven. It drove, it sensed, it responded. It worked.\n\nI am keeping this here precisely because nothing came of it. An archive that only\nholds the projects that continued is a highlight reel, and a highlight reel cannot\nshow you how someone decided what to do.\n\n## The problem\n\nThe stated problem was the usual set of exercises: follow a line, avoid an obstacle,\nrespond to a sensor. The real problem was one I did not know I was running — I was\ntrying to find out which part of building things I actually liked, and I had no\nvocabulary for the question yet.\n\n## The approach\n\nI built the exercises. Line following, obstacle avoidance, a couple of behaviours of\nmy own that were more elaborate than they needed to be.\n\nAnd I noticed, slowly, that the moment the robot moved was the least interesting\nmoment. What held my attention was the part before: getting the logic right, watching\na condition I had reasoned about turn out to be true. The motors were an output\ndevice. I would have been just as satisfied by a printed line of text — which, it\nturns out, is the definition of not being a roboticist.\n\n## Architecture\n\nBlock-based, so there is not much architecture to describe, and that is itself part of\nthe record. Blockly removed syntax as an obstacle and left structure — sequence,\ncondition, loop, state — which is the right thing to meet first.\n\nThe limitation taught me more than the capability. Blockly makes simple things trivial\nand complex things visually unmanageable; a behaviour with four interacting conditions\nbecomes a wall of nested shapes you cannot read. Hitting that wall is what made text\nfeel like a relief rather than a hurdle.\n\n## Results\n\nA working robot and a decision. I stopped doing robotics deliberately rather than by\ndrift, which is the only reason the year counts for anything.\n\n## What I would do differently\n\n- **A dead end you can name is worth more than a success you cannot explain.** I know\n  why I stopped. That answer has redirected everything since.\n- **Enjoying the outcome is not the same as enjoying the work.** The robot moving was\n  the outcome. The logic being right was the work. Those separate early, and noticing\n  which one you are chasing saves years.\n- **Outgrowing a tool is data.** The frustration was not a failure of Blockly. It was\n  the signal that the problems I wanted had passed what it could express.",
      "status": "PUBLISHED",
      "occurred": "2023-10-01T00:00:00.000Z",
      "recorded": "2023-10-01T00:00:00.000Z",
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/projects/mbot-neo"
    },
    {
      "kind": "project",
      "slug": "vector",
      "title": "VECTOR",
      "summary": "An autonomous infrastructure system that monitors itself, contains its own failures, and narrows its own authority when it stops being certain.",
      "body": "VECTOR — Variable Environment Control Through Observation Response — began as a\ntraffic intelligence system and became something with a wider remit: a real-time\nsystem that governs itself through an explicit escalation chain ending at a human\nbeing.\n\nTraffic is still the first application. It is no longer the description.\n\n## The problem\n\nThe original problem was clean. Predict congestion, prioritise emergency vehicles,\noptimise flow in real time.\n\nThe problem that replaced it announced itself as the system grew: **a smart model is\nworthless if it cannot survive failure.** A traffic intelligence platform that goes\ndark during peak congestion is not merely inconvenient — it is dangerous. And the\nmoments that decide whether infrastructure is safe are exactly the moments benchmark\naccuracy says nothing about: when the system is wrong, uncertain, degraded, or being\nfed bad data.\n\nThat reframes what needs building. Not a better decision-maker. A structure inside\nwhich a decision-maker can be trusted to run unattended — bounded, interruptible,\naccountable, and capable of reducing its own authority when conditions stop being\nlegible.\n\n## The approach\n\nTwo layers, built in that order, and the order is the design.\n\n**The governance stack** came first — eight protocols forming one escalation chain.\nSentinel detects anomalies. Serpentine classifies system stress across five levels.\nAegis contains failures before they cascade. Atlas rebalances load. Mender repairs,\ngoverned by a credit budget that makes it progressively more conservative as failures\naccumulate. Phoenix orchestrates recovery. Hyperion is an emergency shutdown that\nfires within a second with no cleverness in it at all. Terminus is the permanent stop,\nand orchestrators, agents and internal services are explicitly blocked from invoking\nit — it requires a human at a physical console.\n\n**The intelligence layer** was allowed on top only once that was proven. Anarchy runs\nbounded what-if simulations on specific triggers. ORACLE predicts futures in latent\nspace rather than simulating them. MORL handles objectives that genuinely conflict.\nPULSAR makes small reversible corrections. Fusion coordinates them under safety gates,\nand Meta adapts the whole thing slowly, and only when conditions are stable enough to\njustify adapting at all.\n\nAbove both sits the authority layer, which escalates deterministically from\nnormal through warning and critical to activation — and on activation strips the\nsystem back to its deterministic baseline before reintroducing intelligence in stages.\n\n## Architecture\n\nThe pipeline:\n\n```\nInput → Sentinel → Serpentine → Anarchy → ORACLE → MORL → PULSAR → Fusion → Meta → Output\n```\n\nDetection and classification run before any intelligence does — the system establishes\nhow stressed it is before deciding how ambitious to be.\n\nFour principles hold it together. **Bounded everything**: every stage has a hard time\ncap and exits rather than running long, because a late decision in infrastructure is a\nfailed decision under another name. **Intelligence is subordinate to safety**: the\ncontrol plane is immutable and no amount of model confidence widens the envelope.\n**Tail risk is first-class**: p95 and p99 are optimised alongside the average, because\nan average latency means nothing if the tail is catastrophic. **Autonomy is earned**:\nunder sustained uncertainty the system narrows its own freedom and restores it only\nagainst evidence.\n\nThe most important structural property is that the deterministic controller built\nbefore any learning is still maintained as a first-class component. It is not\nscaffolding. It is the floor everything falls back to when the authority layer activates, which is\nwhy the system never faces a choice between doing something clever and doing nothing.\n\n## Results\n\nv1 completed a two-hour adversarial stress test across more than 647,000 cycles: zero\nfailures, zero crashes, zero dropped events on the bus, zero ordering violations, and\n6.17 ms p95 under sustained adversarial load. A shorter profiling snapshot measured\n0.607 ms average, 1.124 ms p95, 1.634 ms p99, 1.908 ms maximum, with 32 of 32\nregression tests passing.\n\nThose numbers were bought once, in a single fix. An early profile sat near 100 ms p95\nwith tail spikes that made the system unusable. Profiling found Sentinel — the\nsimplest component, written first and trusted longest — consuming roughly 97% of\nruntime through unbounded statistical work at the head of the pipeline. Rolling-window\nO(1) statistics, a 5 ms hard cap, heavy analysis moved to bounded background tasks, and\nper-stage budget enforcement across the pipeline took it to real-time.\n\nWhat is not done: SUMO training and a real pilot. Until a system has met an\nenvironment it did not anticipate, its numbers describe a laboratory.\n\n## What I would do differently\n\n- **Find the bottleneck before optimising anything.** Six weeks went into components\n  that were not the problem. Suspicion should follow measurement, not architecture —\n  cost does not care how sophisticated a component looks.\n- **Unbounded is the defect, not slow.** Sentinel was fine on average. It was\n  unbounded, so its worst case was unrelated to its typical case, and in a real-time\n  system the worst case is the specification.\n- **Build the floor before the ceiling.** The deterministic controller existed before\n  any model did, which means every learned component has something to beat and every\n  degradation path has somewhere defined to land.\n- **A system that cannot stop is not autonomous, it is unsupervised.** Hyperion and\n  Terminus are the least clever components in the project and the ones I would keep if\n  I had to discard everything else.\n- **The name is load-bearing.** EXTREMIS kept pulling explanations back toward traffic\n  long after the system had outgrown it, and that was shaping the architecture rather\n  than just the conversation.\n- **An accident is not evidence.** Rajasthan is why I returned to this problem and\n  proves nothing about signal timing. Keeping those two statements apart is a\n  discipline I had to learn deliberately.",
      "status": "PUBLISHED",
      "occurred": "2025-10-20T00:00:00.000Z",
      "recorded": "2025-10-20T00:00:00.000Z",
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/projects/vector"
    },
    {
      "kind": "project",
      "slug": "vihaanvaghela-com",
      "title": "vihaanvaghela.com",
      "summary": "A personal archive built to outlast its own framework — CMS, essay pipeline, and an operator console with no URL.",
      "body": "This site. One Next.js application serving three systems that share a database and\nnothing else: the public archive, a CMS I wrote rather than installed, and an operator\nconsole that has no address at all.\n\nI built it because a personal site assembled from someone else's template would have\ndocumented nothing about how I think, which is the only thing it is for.\n\n## The problem\n\nThe requirement that shaped everything: this has to still work in ten years, and it\nhas to be able to show its own history.\n\nThat rules out most of the obvious answers. A hosted platform decides what your\narchive looks like and can withdraw the decision. A static site generator makes\npublishing a deploy, and anything that turns writing into a build step means less\nwriting gets done. A database-backed CMS fixes that and introduces the opposite\nproblem: the About page, which changes twice a decade, now lives behind a login\ninstead of in version control where its diffs are readable.\n\nThe second requirement was awkward. I wanted a VECTOR operator console reachable from\nthis site without being part of this site.\n\n## The approach\n\nContent splits by how often it changes and who needs to read its history. Essays,\nprojects, timeline entries and VECTOR documents are database rows, publishable with no\ndeploy. The About page and site configuration are modules in the repository, because\ntheir diffs are the interesting part.\n\nThe console is reached by typing a trigger into the search bar. It is not linked, not\nin the sitemap, not in robots.txt, and every one of its paths returns a bare 404 to an\nunauthenticated request — byte-identical to a path that has never existed. That is a\nuser-interface property and explicitly not the security model: every credential check\nhappens server-side, and someone who extracts the trigger from the JavaScript bundle\narrives at the same locked door as someone who guessed. Argon2id password, then an\nemailed one-time code, then a WebAuthn passkey.\n\n## Architecture\n\nNext.js 16 App Router on Bun, Prisma over Postgres, Tailwind, deployed on Vercel.\nTwenty-one route handlers, all declaring the Node runtime explicitly, because two\ndependencies are compiled binaries — Argon2 and Prisma's query engine — and no edge\nisolate will load them.\n\nThe provenance split is the decision I am most attached to. `Post` and `VectorDoc` are\nseparate models holding what is superficially the same thing. A post is personal\nwriting and is human-written, which is why that model has no field to record anything\nelse — the absence is the rule. A VECTOR document is AI-assisted and then reviewed,\nand the model records both facts. Publishing one that was never reviewed is refused in\ncode, because a pipeline step nothing enforces is a step that gets skipped.\n\nSecurity headers live in `next.config.ts` rather than a proxy, so they travel with the\napplication. The session cookie is `__Host-` prefixed, HTTP-only and SameSite=strict.\nMarkdown renders without raw HTML, and `dangerouslySetInnerHTML` appears nowhere.\n\n## Results\n\nLive, with every public route served from the database, an RSS feed, a sitemap that\nreflects what is actually published, and a full-text search that the console hides\nbehind.\n\nThe deployment is the part worth recording, because it went badly and the record is\nthe point. Cloudflare Pages could not run the application at all — three hard\nincompatibilities, not a configuration problem. Vercel could, but SQLite could not\nfollow it there. Moving to Postgres broke the container path that had been working and\nleft the documentation describing a database that no longer existed. Then Neon's\npooled connection string turned out to be the wrong one to run migrations through.\n\nNone of that is in the design. All of it is in the git history.\n\n## What I would do differently\n\n- **Ship the thing that makes writing easy.** Every decision that made publishing\n  require a deploy was a decision to write less.\n- **Documentation rots faster than code, and more quietly.** Changing the database\n  provider silently invalidated six statements across two files and a Dockerfile, none\n  of which failed a test. The code complained immediately. The docs told nobody.\n- **Obscurity is a UI property, never a control.** Hiding the console behind a search\n  trigger is good design and worth nothing as security. Writing the second half down is\n  what stops the first half from quietly becoming the plan.\n- **Encode the rule you actually mean.** \"Please review AI-assisted docs before\n  publishing\" is a wish. A publish path that refuses is a rule.",
      "status": "PUBLISHED",
      "occurred": "2026-05-01T00:00:00.000Z",
      "recorded": "2026-05-01T00:00:00.000Z",
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/projects/vihaanvaghela-com"
    },
    {
      "kind": "milestone",
      "slug": "bharuch",
      "title": "Bharuch",
      "summary": "",
      "body": "Born on 15 October 2011, in Bharuch, Gujarat.\n\nI have no memories of the town from this period — everything I know about it arrived\nlater, secondhand, and then in person as a visitor rather than a resident. It is on\nthe timeline because it is the only entry I did not participate in, and because the\ndistance between where a person starts and where they end up is not a detail.\n\nMumbai came much later, by a route that went through New Jersey and Virginia first.",
      "status": "PUBLISHED",
      "occurred": "2011-10-15T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/bharuch"
    },
    {
      "kind": "milestone",
      "slug": "piscataway-new-jersey",
      "title": "Piscataway, New Jersey",
      "summary": "",
      "body": "We moved to the United States when I was four. New Jersey first — an apartment complex\nwith a playground that constituted, as far as I was concerned, the entire world.\n\nMy clearest memory of that year is getting stuck inside a baby swing. Properly stuck.\nAnother child had to work me loose while I considered the possibility that this was\npermanent. I remember the specific indignity of it more sharply than I remember the\napartment.\n\nThe other memory is the one I think about more. My closest friend on that playground\nwas a boy called Harsh, and we spent entire afternoons collecting tiny insects and\nworms and keeping them in bottles overnight. We would come back the next morning to\nfind they had died. I was genuinely upset each time — and then we did it again. And\nagain. I have thought about that a lot since, because it is the earliest version of me\nI recognise: running an experiment, getting a result I did not like, being sincerely\nsad about it, and running it again to see whether the result held. I did not learn the\nlesson the worms were teaching for years. I did learn that I would rather find out\nthan not.",
      "status": "PUBLISHED",
      "occurred": "2015-06-01T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/piscataway-new-jersey"
    },
    {
      "kind": "milestone",
      "slug": "richmond-virginia",
      "title": "Richmond, Virginia",
      "summary": "",
      "body": "We moved south, to Richmond, and I started at Short Pump Elementary School. I turned\nseven there.\n\nVirginia is the single strongest influence on who I am, and I did not notice it\nhappening. It was the first place where I had a real community rather than a\nplayground — friendships that lasted, teachers who argued with me rather than at me,\nand classrooms where being wrong out loud was an ordinary part of the day rather than\na small humiliation.\n\nAlmost everything about how I communicate comes from those years. The willingness to\nsay the half-formed thing and let it get corrected. The assumption that a question is\na contribution. The confidence, which is genuinely a Virginia import and which I would\nnot have grown on my own.",
      "status": "PUBLISHED",
      "occurred": "2016-08-01T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/richmond-virginia"
    },
    {
      "kind": "milestone",
      "slug": "elementary-school-and-no-particular-direction",
      "title": "Elementary school, and no particular direction",
      "summary": "",
      "body": "Four years in which nothing happened that belongs on a résumé, and which I would not\ntrade.\n\nI was not specialising. I was reading widely and badly, running around outside,\nstarting things and abandoning them, and being curious in the unfocused way that only\nreally works before anyone expects an answer from you. Nobody asked me what I wanted\nto be. It was the last stretch of time where the question had not arrived yet.\n\nThis is also where the sport started, unseriously — running because running was\nfaster than walking, and the beginnings of Taekwondo, which was the first thing that\never taught me that improvement can be invisible for months and then arrive all at\nonce.\n\nI am putting a four-year block on this timeline with no accomplishment in it on\npurpose. Growth is not made of milestones. Most of it is made of the years in between\nthem, and a timeline that only shows the peaks tells you nothing about how someone got\nto the next one.",
      "status": "PUBLISHED",
      "occurred": "2017-08-01T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/elementary-school-and-no-particular-direction"
    },
    {
      "kind": "milestone",
      "slug": "scratch",
      "title": "Scratch",
      "summary": "",
      "body": "On 14 September 2020 I opened Scratch for the first time. I was eight.\n\nWhat I remember is not the programming. It is that I could build a world and then be\ninside it. I made games, animations and platformers, and I drew every sprite by hand\nin the built-in editor — badly, at length, with enormous conviction. The most finished\nof them was a platformer called Attack of Excalius, which had a title screen I rewrote\nmore times than the actual game and a boss fight I could never balance.\n\nThe thing worth recording is why it took. It was not the technology. I did not care\nabout the technology and would not have known what to call any of it. It was that\nmaking worlds turned out to be addictive in a way nothing else I had tried was. Every\nother interest at that age was something I did. This was somewhere I went.",
      "status": "PUBLISHED",
      "occurred": "2020-09-14T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/scratch"
    },
    {
      "kind": "milestone",
      "slug": "the-scratch-years-and-then-the-fade",
      "title": "The Scratch years, and then the fade",
      "summary": "",
      "body": "Two years of building things with no purpose beyond building them.\n\nI met people online — other children making other games — and we collaborated on\nprojects, remixed each other's work, and argued about mechanics with total\nseriousness. It was the first time I experienced a technical community, and it was\nmade entirely of eleven-year-olds. I learned more about how to work with other people\nfrom those collaborations than from anything school arranged.\n\nAnd then the interest faded. Not dramatically — it just stopped being the thing I\nreached for. Programming disappeared from my life for a while, and I did not notice it\ngoing.\n\nI am keeping this on the timeline because the tidy version of a builder's history has\nno gaps in it, and mine has one. The interest came back later and came back different.\nIt would be dishonest to draw a straight line through a period where the line simply\nstopped.",
      "status": "PUBLISHED",
      "occurred": "2021-01-01T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/the-scratch-years-and-then-the-fade"
    },
    {
      "kind": "milestone",
      "slug": "selected-for-the-ib-pathway-at-moody-middle-school",
      "title": "Selected for the IB pathway at Moody Middle School",
      "summary": "",
      "body": "During fifth grade I was selected for the International Baccalaureate pathway at Moody\nMiddle School. Academically it was the most significant thing that had happened to me,\nand it is on this timeline mostly for what happened next, which is that I never took\nit up.\n\nI have thought about that pathway more than is reasonable — the version of the next\nfew years that was arranged and then did not happen. It is the first fork on this\ntimeline where I can clearly see the road I did not take.",
      "status": "PUBLISHED",
      "occurred": "2022-11-01T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/selected-for-the-ib-pathway-at-moody-middle-school"
    },
    {
      "kind": "milestone",
      "slug": "back-to-india-into-sixth-grade",
      "title": "Back to India, into sixth grade",
      "summary": "",
      "body": "I left fifth grade half-finished in the United States and returned to India in 2023,\nstarting sixth grade into the ICSE curriculum.\n\nSeven years is long enough to be shaped by a place. It is also short enough that\ncoming back does not feel like coming home — it feels like arriving somewhere you are\nsupposed to already know, with an accent that is wrong in both directions and a set of\nclassroom instincts that do not transfer. I was direct in a way that read as rude. I\nvolunteered answers in a system with different conventions about when you speak. None\nof it was anybody's fault and all of it took a year.\n\nWhat I got out of that year is the thing I now trade on hardest: I stopped assuming\ncontext was shared. When you have lived inside two sets of defaults, you learn that\nthe other person is not being obtuse — they are running a different protocol. Writing\neverything down, which is the habit this entire website is built on, starts here.\n\nThe sport got serious around the same time. Cycling especially, which eventually took\nme to national level, and Taekwondo through to a first-degree black belt. Both were\nhappening in the background of everything below, and both taught the same lesson\nendurance always teaches: impatience is not just unhelpful, it is inert. It does not\ndo anything. The work takes the time it takes.",
      "status": "PUBLISHED",
      "occurred": "2023-06-01T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/back-to-india-into-sixth-grade"
    },
    {
      "kind": "milestone",
      "slug": "blockly-an-mbot-and-the-end-of-robotics",
      "title": "Blockly, an mBot, and the end of robotics",
      "summary": "",
      "body": "Blockly, and an mBot Neo. Line following, obstacle avoidance, a few behaviours of my\nown that were more elaborate than they needed to be. It all worked.\n\nAnd I discovered, slowly, that I did not care. The moment the robot moved was the\nleast interesting moment. What held me was the part before it — getting the logic\nright, watching a condition I had reasoned about turn out to be true. The motors were\nan output device for the part I actually liked. I would have been equally satisfied by\na printed line of text, which is more or less the definition of not being a\nroboticist.\n\nSo I stopped, deliberately rather than by drift, and that is the only reason the year\ncounts for anything. A dead end you can name is worth more than a success you cannot\nexplain. I knew what I was not, which narrowed the search considerably.\n\nBlockly also gave me the other half of it. Block-based programming makes simple things\ntrivial and complex things visually unmanageable — four interacting conditions become\na wall of nested shapes nobody can read. Hitting that ceiling is what made text feel\nlike a relief rather than a hurdle.",
      "status": "PUBLISHED",
      "occurred": "2023-10-01T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/blockly-an-mbot-and-the-end-of-robotics"
    },
    {
      "kind": "milestone",
      "slug": "overheard-engineering",
      "title": "Overheard engineering",
      "summary": "",
      "body": "My father worked remotely, and I did homework near enough to hear it.\n\nGenAI. DevOps. Jira. Atlassian. Retrieval systems. Enterprise software. I understood\nnone of it. What I absorbed was the vocabulary — the shape of the words, the tone\npeople used when they said them, the fact that these were ordinary problems that\nordinary adults argued about on a Tuesday afternoon.\n\nLearning the words a year or two before the concepts is a strange order and I would\nnot design a curriculum around it. But it did something I have come to think is\nundervalued: it removed the intimidation. When I eventually met retrieval systems\nproperly, they were not a forbidding new field. They were a thing I had been\noverhearing since I was twelve. Most of the barrier to entering a technical field is\nthe suspicion that it is not for you, and I had already lost that suspicion by\naccident.",
      "status": "PUBLISHED",
      "occurred": "2024-01-15T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/overheard-engineering"
    },
    {
      "kind": "milestone",
      "slug": "copilot-claude-and-agentic-coding",
      "title": "Copilot, Claude, and agentic coding",
      "summary": "",
      "body": "Modern AI arrived, and programming came back — different from how it left.\n\nThe honest account of what changed is less dramatic than either the enthusiasts or the\ncritics want. It removed most of the typing. It removed none of the thinking. The\nbottleneck moved from \"can I express this\" to \"do I know what I actually want\", and\nthe second question is harder, has no keyboard shortcut, and cannot be delegated,\nbecause it is a question about your own intent and you are the only available source.\n\nThe failure mode I had to learn to guard against is not wrong code — wrong code\nannounces itself. It is plausible code: code that works, that I would not have\nwritten, and that I accepted because it worked. Accept enough of it and you are\nmaintaining a system whose decisions you did not make.\n\nWhich is where the rule came from, and I still hold to it: if I cannot explain every\nline, I do not understand it yet. Not \"could work it out if pressed\" — explain it,\nnow, including why it is this way and not the obvious alternative.",
      "status": "PUBLISHED",
      "occurred": "2024-08-01T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/copilot-claude-and-agentic-coding"
    },
    {
      "kind": "milestone",
      "slug": "rajasthan",
      "title": "Rajasthan",
      "summary": "",
      "body": "We were driving back toward Ahmedabad after a trip through Rajasthan when we were in a\nminor accident. Nobody was seriously hurt.\n\nThe tidy version of this story would say that a traffic system failed and I resolved\nto fix traffic systems. I am not going to write it, because I have no evidence for it.\nI cannot tell you a signal caused that accident. I cannot tell you a better one would\nhave prevented it.\n\nWhat actually happened is smaller. Years earlier, at about ten, standing in New York\nCity, I had watched a traffic light hold a red for a completely empty road and\nwondered why it was running a timetable instead of reading the street. I did nothing\nwith the question — I was ten. In the car, afterwards, it came back, and it came back\nwith an urgency it had not earned on the evidence.\n\nI have decided that is a fine reason to start something and a terrible reason to stop\nchecking your work. Motivation and justification are different objects. The first\nEXTREMIS notes were written within weeks.",
      "status": "PUBLISHED",
      "occurred": "2025-10-20T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/rajasthan"
    },
    {
      "kind": "milestone",
      "slug": "extremis-v1-and-writing-it-down-as-it-happened",
      "title": "EXTREMIS V1, and writing it down as it happened",
      "summary": "",
      "body": "The first real implementation, and — more consequentially — the moment I started\npublishing the reasoning while it was still provisional.\n\nThe early posts are a record of the argument I was having with myself. Why I was\nbuilding systems rather than models. How I reduced the input space to seven variables\nthat stay informative under degradation. Where the boundary between perception and\ndecision should sit, and why collapsing it produces something that reacts fast and\nreasons badly. Why I built a deterministic traffic controller before training anything\nat all, so that every learned component afterwards had a floor to beat.\n\nThen the piece that reframed the whole project: I stopped optimising, and built a way\nto stop. Performance describes a system under favourable conditions. Infrastructure is\nonly interesting under the other ones.\n\nI started writing publicly because I thought universities valued documentation. That\nwas genuinely the reason and I am not going to pretend otherwise. It stopped being the\nreason somewhere in this stretch, and I did not notice the switch until it had already\nhappened.",
      "status": "PUBLISHED",
      "occurred": "2026-01-04T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/extremis-v1-and-writing-it-down-as-it-happened"
    },
    {
      "kind": "milestone",
      "slug": "the-eight-protocol-stack",
      "title": "The eight-protocol stack",
      "summary": "",
      "body": "Sentinel, Serpentine, Aegis, Atlas, Mender, Phoenix, Hyperion, Terminus. One\nescalation chain from the first detected anomaly to a permanent termination that only\na human being can authorise.\n\nThis is the point where EXTREMIS stopped being a traffic model. The realisation behind\nit is a single sentence — a smart model is worthless if it cannot survive failure — and\nonce I took it seriously, the interesting work stopped being the intelligence and\nstarted being the governance around it. A traffic system that goes dark during peak\ncongestion has not underperformed. It has done harm.\n\nTerminus is the part I am most attached to. It blocks orchestrators, AI agents and\ninternal services from invoking it; it requires a person at a physical console with\nmulti-factor authorisation; it encrypts a full state archive before it ends anything.\nThe system cannot reach the end of its own escalation chain. That is not a limitation\nI worked around. It is the feature.",
      "status": "PUBLISHED",
      "occurred": "2026-03-20T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/the-eight-protocol-stack"
    },
    {
      "kind": "milestone",
      "slug": "v1-is-real-647-000-cycles-zero-failures",
      "title": "v1 is real — 647,000 cycles, zero failures",
      "summary": "",
      "body": "A two-hour adversarial stress test across more than 647,000 cycles. Zero failures.\nZero crashes. Zero dropped events on the bus. Zero ordering violations. 6.17 ms p95\nunder sustained adversarial load.\n\nThe number that matters to me is not any of those. It is the one before them: for six\nweeks the system sat at roughly 100 ms p95 with tail spikes that made it unusable, and\nI could not work out why. I profiled everything eventually and found that Sentinel —\nthe anomaly detector, the simplest component, the one I had written first and trusted\nlongest — was consuming about ninety-seven percent of runtime. Unbounded statistical\nwork on every cycle, at the head of the pipeline, poisoning everything downstream.\n\nI did not feel clever when I found it. I felt stupid about the six weeks. Then I felt\nthe thing that keeps me doing this, which is not triumph but quiet — a problem that had\nbeen in the room for a month and a half simply was not there any more.\n\nFind your bottleneck before you optimise anything else. I lost six weeks to searching\nwhere the complexity was, because complexity feels expensive. Cost does not care how\nsophisticated a component looks.",
      "status": "PUBLISHED",
      "occurred": "2026-04-01T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/v1-is-real-647-000-cycles-zero-failures"
    },
    {
      "kind": "milestone",
      "slug": "the-himalayas",
      "title": "The Himalayas",
      "summary": "",
      "body": "I went to the mountains and left the laptop behind.\n\nI climb, and I trail run, and this was the longest I had been away from the system in\nmonths. Mountains and software have the same shape, and it is not the shape people\nusually reach for — it is not about summits. You take a step, then another, and\neventually you arrive somewhere with a view. Then you keep going, because the\nviewpoint was never the destination, only the place where the work becomes briefly\nvisible.\n\nI came back with the clearest head I had had all year. Endurance sport is the only\nthing I have found that teaches patience by making impatience genuinely useless — you\ncannot want your way up a slope any faster. That transfers directly, and it is why I\ndo not treat the training as time taken away from the building.",
      "status": "PUBLISHED",
      "occurred": "2026-05-09T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/the-himalayas"
    },
    {
      "kind": "milestone",
      "slug": "extremis-becomes-vector",
      "title": "EXTREMIS becomes VECTOR",
      "summary": "",
      "body": "Variable Environment Control Through Observation Response.\n\nThe trivial reason is that Extremis belongs, thoroughly, to Iron Man. That alone would\nhave been an annoyance rather than a migration.\n\nThe real reason is that the name had stopped describing the thing and started\nconstraining it. I would sit down to explain the governance stack and find myself\ntalking about intersections — not because Terminus is about intersections, but because\nthe name set the frame and every explanation began at traffic and reached outward\napologetically. That was not a communication problem. Components were being designed\nagainst traffic-shaped assumptions, because traffic was the vocabulary I had.\n\nA name is a constraint on the thinking, not a label on the result. The signal to watch\nfor is not being bored of the name. It is noticing that your explanations of general\nthings keep starting from a specific case.",
      "status": "PUBLISHED",
      "occurred": "2026-06-15T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/extremis-becomes-vector"
    },
    {
      "kind": "milestone",
      "slug": "swift-vector-lite-and-a-physics-exhibition",
      "title": "Swift, VECTOR Lite, and a physics exhibition",
      "summary": "",
      "body": "I picked up Swift and started building native iOS — VECTOR Lite, with an eye on the\nSwift Student Challenge. A second language with genuinely different opinions is worth\nmore than a second framework with the same ones: value semantics, optionals that will\nnot let you pretend a thing might not be there, and a compiler that argues back.\n\nI also took the work to a physics exhibition, where the judges did not understand it.\n\nThat deserves to be recorded accurately rather than bitterly. They were not wrong to\nbe unconvinced. I presented a governance architecture to people expecting a physics\ndemonstration, used vocabulary I had never had to define because everyone I discuss\nthis with online already shares it, and answered the questions I wished they had asked\ninstead of the ones they did.\n\nBuilding something difficult does not entitle you to be understood. Communication is\nnot a thing you do after the engineering — it is part of it, and I had genuinely\nbelieved otherwise until standing in front of a table where the belief did not\nsurvive. It is the most useful bad afternoon I have had.",
      "status": "PUBLISHED",
      "occurred": "2026-07-01T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/swift-vector-lite-and-a-physics-exhibition"
    },
    {
      "kind": "milestone",
      "slug": "vihaanvaghela-com",
      "title": "vihaanvaghela.com",
      "summary": "",
      "body": "The site you are reading. Weeks of work, most of it not the interesting kind.\n\nGitHub, Cloudflare, Vercel, Neon, Prisma, DNS, SSL, Resend. Cloudflare Pages turned\nout to be unable to run the application at all — not a configuration problem, but\nthree hard incompatibilities including a compiled Rust hashing binary and a filesystem\nthat Workers do not have. Then SQLite could not follow the application to a serverless\nhost, so the database moved to Postgres, which broke the container path that had been\nworking and left the documentation describing a database that no longer existed. Then\nthe pooled connection string turned out to be the wrong one to run migrations through.\n\nNone of that is in the design. All of it is in the git history, which is where it\nbelongs.\n\nWhat this actually marks is a change in what I am doing. Before this, I built things\nand wrote about them somewhere else. This is the first time the work and the record of\nthe work live in the same place, under my own name, on infrastructure I control and\ncan still be running in twenty years.\n\nThis timeline is still being written.",
      "status": "PUBLISHED",
      "occurred": "2026-08-02T00:00:00.000Z",
      "recorded": null,
      "assisted": false,
      "supersededBy": null,
      "revisions": [],
      "url": "https://vihaanvaghela.com/timeline/vihaanvaghela-com"
    }
  ]
}