Built for codebases that exceed any single context window by orders of magnitude. Where agents go blind, the engine sees a smooth manifold.
Topaas // Topology-as-a-Service · est. MMXXVI
The Codebase Intelligence Layer between your company and AI.
Laplace ENGINE
Give your large legacy codebase an AI-native soul.
AI coding agents are fast on small tasks. They break on large, legacy systems. Laplace turns large, complex, legacy codebases into agent-ready system intelligence — so AI agents can reason over them with full context.
ℒ { codebase } ⟶ Lv = λv ⟶ spec(L)
01 The Shape of Code
Coding agents see tokens.
Engineers see files.
A codebase is a manifold.
Functions, modules, and data flows form a high-dimensional simplicial complex with measurable topology — homology, fundamental group, spectrum. Most of what makes legacy systems hard is invariant under refactoring. We compute the invariants.
Ingest
Raw code · symbols · runtime traces · commit history
C = ⋃i fi
Lift to a graph
Build the dependency complex G = (V, E) and weight edges by behavioural coupling.
G ↪ K(C)
Compute the Laplacian
Diagonalise L = D − A. The spectrum encodes clusters, bottlenecks, and the true seams of the system.
L vk = λk vk
Hand to the agent
The spectrum becomes the spec — components, intent, behavioural commitments, risk surfaces: the build order an agent can execute, at a scale no context window holds.
spec(L) ≡ the spec
spec(L) reads two ways — the spectrum of the Laplacian, and the specification your agent builds from. That's the whole idea.
02 Field Operators
Three invariants we hold quietly, for now.
We extract the spectrum of the Laplacian — clusters, connectivity, bottlenecks, components. The unit of work becomes the eigenmode, not the line.
Software deforms over time. We track the system up to homotopy — keeping the topology coherent as code, intent, and behaviour drift.
The same operator — Laplace's — runs through signal processing, graph theory, Riemannian geometry, and quantum mechanics. We brought it to legacy software.
03 Field Readings
We don't assert the spectrum.
We measure it.
Early readings from the engine, run against open substrates at production scale — codebases an order of magnitude beyond any context window. Numbers are outcomes only: answer quality held at parity, the implementation held quiet.
Fewer tokens, same answer.
≈ 590K ⟶ 11K tokens / query
Find the right file, top-5.
issue ⟶ the right file · file Acc@5
Correctness on a full rebuild.
apache/hadoop · full reconstruction
Lines lifted into one graph.
~117K nodes · ~419K edges
Measured on open-source substrates at production scale; the full rebuild ran on Apache Hadoop. Baselines are full-context and single-shot agents; answer quality is held at parity by an independent judge. Other substrate identities and our production pipeline stay folded while we're in stealth.
04 Provenance
A small group, quietly assembled.
Researchers and engineers who have built large-scale ML, data, and distributed systems at the kind of places where scale is the default. Names stay folded while we're in stealth — open a transmission and we'll introduce ourselves.
- AWS AI
- Meta
- Microsoft
- Wharton Research Data Services
Past employers of individual team members. For identification only — no affiliation, sponsorship, or endorsement is implied.
05 Transmit
Quiet, not closed.
Investors, prospective design partners, engineers who think about software as a topological object — write to us.