The accelerator is a phase.
The processor is the destination.
Agentic AI is not a model you call. It is a program that runs. A4 is the general-purpose, massively parallel processor built for that program: hundreds of independent agents on one die, bound into one system by a nervous system in silicon.
The Processor Learns to Think
Why agentic AI belongs on a general-purpose processor, and how we got there first.
Every accelerator that mattered came home to the processor.
A new workload appears, important enough for dedicated silicon but too immature for the processor to absorb. It ships as a coprocessor. Then it matures, the operations standardise, and the function folds into the main processor, where the data already lives and the rest of the program already runs.
| Era | Function | Started as | Came home as |
|---|---|---|---|
| 1980s | Floating point | Intel 8087 / Weitek coprocessors | On-die FPU, 486DX |
| 1980s–90s | Memory management | External MMU chips | Integrated MMU / TLB |
| 1990s | Vector / DSP math | Discrete DSPs, add-in boards | MMX, SSE, AVX, NEON |
| 1990s–2000s | Graphics | Discrete graphics cards | Integrated graphics / APU |
| 2000s | Cryptography | Crypto accelerator cards | AES-NI, SHA extensions |
| 2010s | Network offload | TCP-offload NICs | On-core packet processing |
| 2020s | AI / neural | Discrete NPU / TPU / GPU | We are here |
We offer this as a track record, not an analogy, and we have skin in it. The architect behind A4 designed CP8, one of the first commercial fine-grained SIMD engines, and commercial floating-point silicon during the very transition when floating point left the coprocessor and joined the CPU. We did not read about the last subsumption. We shipped through it.
AI is now showing every signal the FPU gave in the late 1980s: operations consolidating, precisions stabilising, deployment moving from the datacenter to the edge. The subsumption of AI into the processor is not a prediction. It is a schedule. A4 is built to be early on it.
An agent is a control loop with a neural network inside it. Inference is one phase of seven.
A4 corrects this by refusing the split. The whole agent, control, retrieval, crypto, signal processing and inference, runs on one substrate. No phase of the loop has to leave the chip to find the hardware it needs.
Not a savant. A journeyman, and we mean it as the highest compliment the market can pay a processor.
- Does one thing at the edge of the possible
- Dense matmul, magnificently
- Inert at everything an agent actually spends its time on
The neural accelerator.
- Broadly skilled, competent across the whole trade
- Genuinely excellent at the specialty the work demands most
- Finishes the job. The whole job.
A4.
The physical embodiment is a multi-precision datapath. One engine spans the precisions an agent needs, from binary networks to floating-point signal work, rather than a stack of fixed-function blocks that idle whenever the workload steps outside their class. How A4 sustains that range without the usual efficiency penalty is the heart of the design, and it is available under NDA.
Breadth is only a virtue if it is real. A4's is structural.
Multi-precision inference across a fleet of independent agents, each running its own model at its own precision.
Native floating point for sensor front-ends, filtering and the scientific math neuromorphic and NPU parts simply lack.
On-die hardware root of trust and symmetric / hash acceleration, plus bit-sliced parallel ciphering across the SIMD fabric. The authenticate phase never leaves the chip.
Hundreds of independent, hardware-isolated processors, each running the branch-rich orchestration that is the agent's actual spine.
Because all four live on one substrate, the entire agent lives on one substrate. The two-chip tax is not reduced. It is eliminated, because there is no second chip.
A generic translator leaves most of a journeyman's performance on the floor. So we replaced it.
The compiler is built to target a machine it was not co-designed with, to lower any program without understanding the intent behind it, and to be correct on the average case at the cost of excellence on the specific one. In its place, A4 uses an LLM-driven distiller.
Application code, or a model in a portable form. Standard front doors, including WebAssembly and ONNX, so existing code flows in without being rewritten.
The distiller descends to the metal the way an expert would, with knowledge of the machine a generic compiler never has. Compact, hand-quality kernels.
We do not trust the LLM. We verify it. Every kernel is checked for equivalence against its reference before it is ever allowed to run. Mistakes die in verification, not in silicon.
Each kernel joins a growing library of proven primitives. The platform gets faster and broader with every workload it meets.
Instead of a fixed compiler that ages against a fixed ISA, A4 has a distiller that improves, with a verification floor that keeps every step honest. Generality gave us the compiler. Specialisation, made safe by verification, lets us replace it.
The Nervous System, Not the Muscle
Why agentic AI scales sideways, and the coordination substrate the GPU era forgot.
Build a bigger engine, or build more engines and connect them. The whole history of the field is this argument.
- Deeper pipeline, wider datapath, larger dense array
- Faster and fatter memory feeding it
- The vector supercomputer, the datacenter GPU and NPU
It won the last thirty years, and it won honestly. Dense linear algebra and model training are vertically shaped.
- Many modest processors, each doing a share
- The dataflow machine, the Transputer, the Connection Machine
- The neuromorphic parts that fire on events, not clocks
It lost, not because the idea was wrong but because its workload had not arrived. Every lateral machine died on the same reef: paying for communication the job did not need.
The claim is not that the losers were secretly right. It is that the workload finally changed, and the axis that kept losing is the axis on which agentic AI is won.
A fleet of agents is not a bigger kernel. It is a distributed system, and the scarce resource is coordination.
You do not make a distributed system faster by buying a faster CPU. You make it faster by making the coordination cheap.
Evolution never built the one enormous neuron.
It built many modest ones and spent its real innovation on two things that have nothing to do with making a single cell more powerful.
These are precisely the two properties a conventional compute array lacks. A4 adopts both as architecture, not as software running on top of hardware that fights them. This is a claim about coordination and energy, not about learning, and we hold to that distinction.
A society does not have a foreman issuing every instruction. Its members coordinate directly.
Communication as a first-class silicon primitive is what turns hundreds of processors into a system of agents rather than a pool of workers. How the fabric routes, arbitrates and coordinates is available under NDA.
A machine built to hold a dense array at peak does not gracefully spend less when the work arrives as sparks.
You are, much of the time, paying to keep a furnace warm. A4 does what the nervous system does: spends energy where agents are working and withdraws it from those that are not. Because the fabric already knows where the activity is, who is publishing, who is waiting, the machine has exactly the signal it needs.
Vertical scaling's efficiency story is told at peak, on the one workload that keeps the muscle full. The lateral story is told in the far more common case, the machine half-quiet. On a spiky, many-agent workload, the machine whose energy follows activity wins the metric that actually gets paid: work per watt in the real duty cycle, not at an unreachable peak.
Lateral and vertical are not enemies.
Every agent, at the bottom of its loop, still runs a neural network, and that is dense arithmetic, the vertical machine's home ground. A superb nervous system wired to feeble muscle is a coordinated way to run models too slowly. The lateral bet does not dispense with per-node compute. It governs it.
Strong journeyman nodes, governed by a lateral nervous system. The muscle stays necessary. The nervous system stops being optional.The Nervous System, Not the Muscle, 2026
Decades of silicon-proven fine-grained SIMD. The same hand, three generations.
- ICL DAP pioneers processing-in-memory. Not ours: the lineage A4 descends from.
- UK System-X telephony processor + RTOS. 20,000 processors only now being replaced, 40+ years on.
- AMT, spun out of ICL to commercialise the DAP SIMD line. DAP 500 and DAP 600.
- CP8 coprocessor. One of the first commercial fine-grained SIMD engines. Powered the CPP Gamma series through the 1990s in radar, signal processing and defence.
- Intel-compatible processors. Three generations of x86-compatible designs: two in 1992–98, including the floating-point unit as floating point moved from coprocessor into the CPU, and a third, with SIMD, in 2005.
- Custom-processor IP, developed and licensed. Continuous since 1998, alongside engineering-practice consulting to silicon startups.
- DARPA-funded cognitive-sensor programme. Architect and IP supplier for its multi-TeraOP synaptic processor.
- First-generation prototype fabricated, first parts now back from the fab.
- A4. The agent-native processor. Its compute lane is the direct successor to CP8.
Accelerators always come home. An agent is a program, not a kernel. The journeyman processor. One substrate, many trades. We removed the compiler.
Public edition v1.0 · July 2026Two ways to scale. The workload changed shape. Intelligence is a nervous system. Publish, subscribe, coordinate. Power that follows activity.
Public edition v1.0 · August 2026Both argue the what and the why, not the how. Mechanism-level detail, specifications and quantified performance are withheld and available under NDA. Statements about future products are forward-looking.
History says AI comes home to the processor. Agentic AI says the processor must be broad. Agentic scale says it must be a nervous system. A4 is the bet that all three are true at once.
The public editions of both papers are available on request. The detailed platform architecture and measured results are disclosed under NDA.