Abstract
An AI company is modeled as an open, resource-dependent system whose observations, learning procedures, deployment actions, and governance rules close several feedback loops. A contextual -topos supplies a language for states, admissibility certificates, and coherent identification of implementations; categories and lenses retain the direction of irreversible operations. We give elementary results on certified update preservation, descent, the limits of observational regulation, and the insufficiency of semantic equivalence for identifying learning dynamics. At the differentiable level, an optimizer-dependent neural tangent kernel is derived exactly along gradient flow. Parameter jets, a neural tangent hierarchy, and an input-jet tangent kernel describe distinct extensions. Their order is independent of homotopy truncation. Every extensional optimizer on smooth objectives factors through their global holonomic jet representation, with runtime and oracle state retained; factorization through a jet at one point has a separate locality condition. A formal panoptic hypothesis specifies trace coverage and inequalities between observation, intervention, and governance capabilities. The construction combines identity types, univalence, and jets into a research program for the cybernetics of AI firms, without treating economic concentration, universal surveillance, or a fixed-kernel description of every neural system as mathematical consequences.
Keywords. Organizational cybernetics; homotopy type theory; -topos; lenses; neural tangent kernel; jets; constrained updates.
The company-level application is a proposed synthesis. Its starting hypotheses are that employee and customer traces can both become productive training data, that feedback changes the behavior being measured, and that data integration can create incentives for consolidation. These are conditional empirical and institutional claims. The mathematical results below concern explicitly specified models; they do not establish that every actual AI company collects all available traces or that a single company must emerge.
Wiener’s control-and-communication perspective [35], Ashby’s requisite variety [1], and Beer’s organizational cybernetics [2] motivate the separation of operation, coordination, adaptation, and policy. Modern categorical cybernetics supplies compositional interfaces and controllers [4, 27]; categorical learning supplies parametrized maps, learners, and reverse differentiation [9, 7]. Higher groupoids add a precise account of witnessed equivalence and its coherence. They are not substitutes for economic evidence or for differential equations.
Two further strands connect architecture and institutions: categorical deep learning studies algebraic architecture constraints [13, 12], while compositional game theory models interacting decision-makers [14]. Hedges explicitly proposes categorical-cybernetic analysis of AI-mediated markets and supply chains [16]. That proposal is a research agenda, not an established empirical law of AI companies.
Central question. Which observations and interventions can a firm perform, which specifications must those interventions preserve, and how do learning, governance, and the environment respond to one another? A model is useful when it makes these questions executable or falsifiable, rather than merely relabeling a company as a category.
A category has objects, arrows with specified source and target, identity arrows, and associative composition. A functor preserves these data. A natural transformation consists of arrows between two functors’ values which commute with every source-category arrow. “Small” means that the objects and arrows belong to the chosen set-sized universe. reverses arrows, and denotes functors and their natural transformations. A groupoid is a category whose arrows are invertible. An endomorphism has equal source and target; an automorphism is an invertible endomorphism. A terminal object has one arrow from each object, and an initial object has one arrow to each object, with the corresponding mapping spaces contractible in higher categories. Products and limits are defined by their universal mapping properties.
For a precise higher model, a simplicial set is a functor , where has finite ordered sets and order-preserving maps. The horn consists of all faces except the th. An -category can be modeled by a simplicial set filling every inner horn (); a Kan complex fills all horns and models an -groupoid. The -category of spaces means homotopy types, modeled by Kan complexes with weak homotopy equivalences inverted. A weak homotopy equivalence induces a bijection of components and isomorphisms of all homotopy groups. The group consists of based homotopy classes of maps from the -sphere to , for . All higher limits and pullbacks below are homotopy-coherent limits [25]. The core of a category retains its objects and equivalences.
In the internal type language, a universe classifies the chosen small types. A dependent family assigns a type to each . Its dependent sum contains pairs ; its dependent product contains sections assigning such a to every . A witness is an inhabitant of the stated type. A type is contractible when there exist and paths for every . For , its fiber at is . An equivalence is a map with contractible fibers; is the type of these maps and witnesses. A homotopy from to is a section of . Intensional equality uses such identity types, distinct from the definitional equality used in evaluation rules. The type consists of mere propositions, defined below; has two alternatives. Decidability of a proposition means a witness of , not just the formation of its type [32].
A presheaf of spaces is a functor . A site is a small category with a Grothendieck topology: specified covering sieves, where a sieve is a collection of arrows to an object closed under precomposition. The axioms require the maximal sieve to cover, stability under pullback, and transitivity of covering. A sheaf has compatible local sections which glue uniquely, or up to coherent homotopy in the space-valued case. Descent data include sections on a cover and agreements on all iterated overlaps. The Čech nerve lists these overlaps; its limit is the space of such coherent data. An -topos is an accessible left-exact localization of a presheaf -category. Here a localization has a fully faithful right adjoint; left exact means preserving finite limits, and accessible means preserving -filtered colimits for some regular cardinal . A diagram is -filtered when every subdiagram with fewer than arrows has a compatible cocone. A geometric morphism is an adjoint pair with left exact [25, 22].
An open system exchanges inputs and outputs across a chosen boundary. A controller chooses inputs from observations and possibly memory; feedback feeds outputs into subsequent inputs. Governance specifies who may change controllers, objectives, interfaces, and admissibility rules. Resources are stocks or capacities consumed or replenished by operations. These terms acquire explicit maps, state variables, and authorization relations in the constructions below.
Work in intensional dependent type theory, with a chosen univalent universe when univalence is used. For , an identity witness has type ; for there is a further type , and the process iterates. Reflexivity and identity elimination induce inverses and composition together with higher coherence. Identity elimination says that a family indexed by has a section for all once its reflexive instances have sections, with the corresponding reflexive computation rule. It is the dependent substitution rule used for equality witnesses. This identity tower admits a weak -groupoid structure [24, 33]. It is stronger information than a binary relation of indistinguishability.
The groupoid model demonstrates that uniqueness of identity proofs is not a general consequence of intensional type theory [17]. Martin-Löf’s foundational presentation [26] and Voevodsky’s homotopy -calculus notes [34] provide foundational background. Grothendieck’s homotopical perspective [15] connects suitable models of higher groupoids to homotopy types.
Definition 2.1
Truncation levels
A type is a mere proposition, or -type, when any two inhabitants are equal. It is a set, or -type, when each of its identity types is a mere proposition. Inductively, it is an -type when its identity types are -types. The reflection into -types is written .
Its universal property is for every -type ; denotes the mapping space. The -types are the contractible types. Calling a proposition “yes or no” does not make it a decidable Boolean: is an additional assertion in constructive logic. Likewise, “discrete” in this note means homotopy -truncated; it need not mean discrete as a topological or smooth space. These distinctions matter when software must produce an actual certificate.
Definition 2.2
Contextual firm semantics
Choose a small site . Objects of describe contexts such as a team, project, deployment, jurisdiction, or accessible interface; arrows describe restriction of context. The topology specifies which compatible local views jointly cover a context. Set
Firm states, interfaces, specifications, and equivalence witnesses are objects or structured objects internal to .
This is an actual -topos, obtained by accessible left-exact sheafification [25]. The selected site is a presentation of the semantics. Neither the bare organization nor an arbitrary category of neural networks automatically forms a topos. A -truncated part supports ordinary sheaf logic and the Mitchell–Bénabou language [22]; suitable universes and dependent type-theoretic structure support the higher identity language [32]. An arbitrary such model does not automatically supply every desired computational rule for a proof assistant.
A state in context is a generalized element . A global state is . A point of a topos, a geometric morphism in the ordinary setting, is a different notion. Local inhabitation, a global section, and an executable construction must not be conflated: a semantic witness becomes an effective procedure only through an appropriate constructive interpretation or implementation.
For differentiation, choose smooth parameter spaces and derivative operations. A smooth manifold is a Hausdorff, second-countable space covered by Euclidean charts with smooth transition maps; smooth maps are smooth in those charts. A concrete option is to replace the context site by a product with a site of Cartesian spaces and smooth maps, with the corresponding product topology, and to use smooth finite-dimensional charts for parameter objects. Probability is extra structure too. A standard Borel space is measurably isomorphic to the Borel space of a complete separable metric space. A Markov kernel assigns a probability measure to each , with measurable for every measurable . Composition is . A Markov category packages such composition in a symmetric monoidal category, with compatible commutative copy maps and delete maps , the latter natural and the unit terminal [11]. The tensor denotes parallel composition; copy need not be natural for stochastic maps. Neither calculus nor probability follows merely from the word “topos.”
Remark 2.3
Three independent resolutions
Homotopy order , differential jet order , and access/capability resolution answer different questions. Here labels an observation-and-action interface, not a numerical dimension; it is unrelated to the cardinal used in the accessibility convention. A representable smooth sheaf is the functor for a smooth space . A smooth representable parameter space can be a -truncated sheaf while possessing derivatives of all orders. A finite groupoid can have nontrivial identity structure without possessing any calculus. Increasing bandwidth or acquiring a vehicle changes accessible observations and actions; it does not by itself increase a homotopy truncation level.
Let the ambient state object be
where describes human roles and capabilities, data and provenance, models and parameters, optimizer/runtime memory, controllers, resources, and governance. Products may be replaced by dependent sums when, for example, an optimizer’s state depends on the architecture. Specify a predicate and define
Thus an inhabitant of contains both a proposed state and its validity witness. The predicate may include schema equations, authorization, resource bounds, deployment tests, and declared contractual conditions. Their concrete meaning must be supplied; a proof of the chosen predicate is not proof of every desirable social outcome.
Definition 3.1
Certified update
For an admissible input family , a certified update is a procedure which, given , , and , returns . An optional transition relation can be included in the certificate. When a total implementation branches on acceptance, it must supply a decision procedure for its chosen checks; arbitrary semantic consistency is not thereby decidable.
Proposition 3.2
Preservation under iteration
Suppose an initial state has a witness and every accepted update is certified. Then every state reached by a finite sequence of accepted updates has a witness of . A rejected update that returns the original certified state also preserves .
Proof. The base case is . At each successor step the update supplies the next witness. Rejection reuses the existing witness. Induction on the sequence length gives the claim. No uniqueness of certificates or excluded middle is required.
This formalizes constrained database updates and speculative changes. A valid rollback concerns the modeled state: deletion of already transmitted information, or reversal of an irreversible physical act, requires additional operations and cannot be assumed from this proposition.
For an ordinary categorical database schema , an instance is a functor [31]. Functoriality enforces the schema’s declared path equations. Authorization, statistical validity, and cross-table conditions beyond those equations must be specified separately. For a schema map , restriction is precomposition by ; when they exist, its left and right adjoints and are the left and right Kan extensions. An adjunction is a natural equivalence between the corresponding mapping spaces. These migration constructions are not automatic proofs of the additional conditions.
Proposition 3.3
Descent of certified operations
Let be a cover. Suppose local states, transition maps, and certificates form a homotopy-coherent descent datum in their respective sheaf objects. Then they determine a certified operation over , unique up to the contractible space prescribed by descent for that datum.
Proof. Apply the sheaf equivalence between sections over and the limit of sections over the Čech nerve to the structured object of certified operations. This object is assembled from dependent sums, products, and the specified certificate family. The fiber over the given descent datum is contractible.
Independent local approvals are not enough. Agreement on shared data, interfaces, and higher coherence is part of the premise. For example, two resource requests may each fit a budget locally and exceed it jointly; they do not form a valid global resource certificate.
A deterministic open module consists of a state object , an output map , and an update . Closing an environment/controller map gives
More generally, has its own state and the environment contributes inputs; the full product state must then be retained. In a stochastic interpretation, updates are Markov kernels and composition uses their integral composition law.
For interfaces and , a simple lens is a pair
The second component carries responses, requested changes, or cotangent signals back through the interface. For and , composition is
This is the elementary bidirectional structure; dependent lenses permit response types to depend on the current output. Reverse differentiation sends a smooth map to forward evaluation and cotangent pullback in Euclidean coordinates. The chain rule makes this construction compositional. An optimizer closes the parameter port, as in categorical learning [7]. The tangent space at a smooth point consists of derivations on smooth germs, linear over and satisfying ; these are infinitesimal directions. Its dual, the cotangent space, contains linear functionals on those directions. A parameter port is the part of the interface carrying parameters and their requested changes. To close it is to supply an update law for those changes.
Parametric maps and architecture constraints. A layer with parameters is a map . Composing it with gives a map with parameters , . A reparametrization relates to when . Such maps form the elementary structure behind the bicategorical construction; invertible reparametrizations give only its equivalence part. A -category has objects, -arrows, and -arrows between -arrows, with compatible vertical and horizontal composition; a bicategory weakens associative and unit laws to coherent invertible -arrows. An optic represents bidirectional behavior through a residual object , with maps and . Representatives are identified by the relations induced by changing residual objects through compatible maps. Thus residual state is retained for the backward response [12].
Algebraic architecture constraints can also be expressed through monads. In an ordinary category, a monad consists of an endofunctor and natural transformations and satisfying
An algebra satisfies and . These equations constrain how composed operations are implemented. Categorical deep learning lifts this algebraic viewpoint to appropriate -categories of parametric maps, relating model constraints to architecture implementations [13]. The lift and the chosen ambient category are substantive parts of the construction. Architecture algebra does not by itself specify loss, optimizer memory, or a firm’s governance objective.
The temporal operations in (2) are generally irreversible. They belong to a category of transitions or systems. Only equivalences belong to its maximal higher groupoid. In an -category higher morphisms are invertible; its ordinary -morphisms need not be. Reflexivity, symmetry, and transitivity of a relation alone do not provide the entire identity tower.
An executable trace with specified state semantics denotes the composite when the types match. This makes a repeatable operation precise. An untyped log or a trace with missing causal context does not canonically denote an endomorphism. Recursive improvement can be modeled by an endomorphism of a type of certified agents; iteration additionally requires preserved admissibility and realizable operations, not merely a self-description.
Univalence asserts that the canonical map
is an equivalence [32]. It permits transport of type-dependent properties and structures along a supplied equivalence. The identity structure includes automorphisms; univalence does not reduce all equivalence witnesses to one proof. Nor is it derivable just from a verbal appeal to Leibniz’s principle. The structuralist interpretation motivates a chosen foundation; it does not establish a uniqueness theorem for all possible foundations.
For firms or learners, equivalence must be defined on structured systems. An appropriate object includes ; its equivalences preserve the declared interfaces and these chosen structures. Equivalence of underlying state types alone transports no unspecified costs or control policies. A chosen universe may forget them entirely.
Example 4.1
Equal predictions, different learning speeds
Let and , both mapping to . The reparametrization is a smooth bijection and . For the same output loss and Euclidean gradient flow in each coordinate, the scalar tangent kernels are and . Hence
The models have the same realizable predictions and equivalent parameter spaces, but different time-dependent predictions from matched initial outputs.
Proposition 4.2
Kernel invariance requires optimizer geometry
Suppose for a smooth diffeomorphism , and let and be positive semidefinite cotangent-to-tangent preconditioners. If , then the output kernels agree at matched parameters.
This gives a concrete requirement for a higher groupoid of learner presentations: equivalences of optimizer-equipped systems must preserve the relevant geometry. Gauge symmetries such as hidden-unit permutations may be organized into an action groupoid, and, with descent, a quotient stack . Its stabilizers record presentation automorphisms. A quotient is not automatically a smooth manifold, and a kernel descends only when the required invariance holds. Here a group action is a map preserving group multiplication and identity; its action groupoid has an arrow for each . The stabilizer of is . A stack is a groupoid-valued sheaf satisfying coherent descent, and the quotient stack is the sheafified action groupoid. Gauge symmetry here means a reparametrization preserving the declared learner structure. A moduli object classifies such structures together with their equivalences; it retains automorphisms rather than only their orbit set.
Example 4.3
Truncation does not identify realizations
For every , the point and have equivalent -truncations but are not equivalent spaces, since the latter has . Even full behavioral equivalence does not fix a machine: adjoining unused memory to a deterministic implementation leaves its observed behavior unchanged while changing its physical state count and possible cost.
The point is the one-point homotopy type. The Eilenberg–Mac Lane space , for , is connected, has , and has all other positive-degree homotopy groups zero. Behavioral equivalence means equality of the declared observable input-output behavior; additional internal state is retained only when it is part of that declared structure. Thus a unique specified operation class need not have a unique implementation. Also, is not in general a left-exact localization of an -topos. In spaces,
so -truncation fails to preserve this homotopy pullback. Sheafification and homotopy truncation must therefore play distinct roles.
The notation means derivatives through order exist and are continuous; means this for every finite . Write for the Jacobian, for the -linear derivative, for the gradient in the declared inner product, and for its Hessian. A dot denotes a time derivative. A real symmetric matrix is positive semidefinite, written , when for every ; means . A Gram matrix has entries given by inner products of chosen vectors and is positive semidefinite. Neural tangent kernels below are matrix-valued Gram kernels; their meaning differs from a probabilistic Markov kernel. A loss is a scalar objective measuring the chosen discrepancy. A preconditioner maps its parameter covectors to update directions.
Fix a dataset and concatenate a differentiable network’s training outputs into . For vector-valued outputs includes both sample and output indices. Let be differentiable, set and , and prescribe
Here is the number of parameter coordinates, the number of stacked output coordinates, and the sample count when a sample-normalized loss is used. A neural network means the chosen parametrized map ; stacks its outputs on that sample. The transpose is denoted , is the identity matrix of the required size, and is the Euclidean norm or its induced multilinear operator norm. The matrix is a chosen preconditioner and is an additional parameter drift. In geometric language maps covectors to tangent vectors; the transpose uses the displayed Euclidean coordinates. Assume sufficient local regularity for the trajectory under consideration to exist. The following is a direct chain-rule calculation, not an infinite-width approximation.
Theorem 5.1
Optimizer-dependent tangent dynamics
For , , this is the empirical neural tangent kernel (NTK). In sample notation, including the output block structure,
Jacot, Gabriel, and Hongler established the foundational NTK viewpoint and a constant-kernel infinite-width regime under their hypotheses [19]. Linearized wide-network results [23] and tensor-program calculations for broad architectural classes [36] extend its scope. They do not imply that an arbitrary finite transformer or other deployed network has a constant kernel. Feature learning can change and ; lazy training is a particular scaling regime [5] in which the parametrized map remains close to its linearization at initialization over the specified training interval.
For , sufficiently differentiable and no explicit time dependence of ,
The evolution already involves second parameter derivatives, and derivatives of the optimizer geometry. For the special scalar-output Euclidean flow with , define recursively
The labels are fixed scalar targets, , and in is an ordered -tuple of sample indices; this local use is distinct from the identity matrix notation. Whenever the derivatives exist, the chain rule yields the exact identities
This is the recursive structure of the neural tangent hierarchy (NTH) [18]. These higher tensors are not, in general, positive semidefinite two-input kernels or fully symmetric tensors. Freezing a finite level is a closure approximation; the identities themselves do not bound its error. Huang and Yau obtain controlled truncations under specified smoothness, width, initialization, and data assumptions. Their approximation theorem is not asserted here for every modern architecture or optimizer.
Fix an open parameter domain ; the same definitions apply in manifold charts. For scalar objectives write . The standard geometric constructions of jets and their coordinate transformations are developed in [30] and [21]. The factorization results below are elementary consequences of explicitly retaining zeroth-order values and runtime state.
Definition 6.1
Smooth germ and pointwise jet
The germ of a smooth function at is its equivalence class under agreement on some open neighborhood of . The germ algebra is , where the maps are restrictions and the colimit runs over neighborhoods . Two germs have the same order- jet when all their derivatives of orders at most agree at . Their infinite jets agree when this holds at every finite order. The infinite jet fiber is the inverse limit
with bonding maps forgetting the highest-order derivatives.
An inverse limit here is a compatible sequence under these forgetful maps; the germ direct limit identifies representatives agreeing after restriction. In Taylor coordinates the natural map sends a germ to its full derivative sequence. Its kernel is the ideal of flat germs, whose derivatives of every order vanish at . It can identify distinct smooth germs. In one variable, for and for has a nonzero germ at zero but zero infinite jet there. Every derivative of on is an inverse-power polynomial times the exponential and tends to zero at the origin, which proves flatness. An analytic germ is represented by a locally convergent power series; for analytic germs the Taylor map is injective. Smooth germs, formal infinite jets, and analytic germs therefore require distinct types.
Definition 6.2
Global holonomic jet representation
The prolongation of is its jet section over the entire domain, . A section is holonomic when it is the prolongation of an actual smooth function. Let be precisely these sections. Zeroth-order projection defines by . Thus and . The global section retains at every point. This differs from the derivative sequence at a single point. With unnormalized coordinates , holonomicity implies , where is the th unit multiindex. Conversely, smooth component functions satisfying these equations are the derivatives of , by induction. Arbitrary pointwise formal coefficient assignments need not satisfy this compatibility.
The assignment is a sheaf: local smooth functions agreeing on overlaps glue to one smooth function. Germs are its stalks, namely the direct limits over neighborhoods. Prolongation and zeroth-order projection commute with restrictions on the corresponding holonomic jet sheaf. This gives the local-to-global organization of objective data inside smooth contextual semantics. An algorithm querying outside an open set is not an operation determined solely by the objective restricted to .
In a chosen parameter chart, the order- parameter jet of a map is its Taylor polynomial modulo terms of degree :
At a fixed point it records derivatives through order , not the entire map. Under coordinate changes these records transform by the chain rule. They can be organized into jet bundles, whose fibers collect the jets at each base point, and compositional differential structures [3, 30, 21]. Here is a multiindex, , , and . For a one-variable jet over a characteristic-zero algebra , normalized Taylor coefficients obey
Equivalently, with unnormalized derivatives, . Thus multiplication is a convolution, not pointwise multiplication of derivative lists. The product order shown also works for noncommutative coefficients when the variable is central and is a derivation; arbitrary reordering is not permitted. An algebra here is an associative algebra over a field; characteristic zero ensures the displayed integer coefficients have their usual nonzero values.
Lemma 6.3
Differentiation lowers finite jet order
For , differentiation of germs induces , but ordinary polynomial differentiation does not induce an endomorphism of for a characteristic-zero field .
Proof. If two representatives differ by , their derivatives differ by a multiple of , so they have the same -jet. But is zero in the quotient and its derivative is nonzero there. An endomorphism defined by differentiating representatives is therefore not well defined.
The infinite jet, or formal power series, admits the coefficient shift without this finite-order loss. Truncated polynomial representatives can be differentiated after choosing a representative, but that is a chosen approximation to the germ’s derivative. Regular jets do not automatically include Laurent series with a pole at the base point. Shared differential formulas are not an equivalence of all polynomial, Laurent, smooth, and formal theories.
Definition 6.4
Stateful objective-based optimizer
Let be a runtime state space containing current parameters and any memory, query history, population of candidates, or random seed used by the algorithm. Let contain declared auxiliary data, constraints, precision, schedules, objective presentations, and oracle access rules. A deterministic optimizer step is a possibly partial map , where contains subsequent states or stopping results. For a single-state iteration one can take , where is the declared set of terminal answers and statuses. It is extensional in the objective when equal functions with the same runtime state and auxiliary data yield the same definedness and output. A stochastic step is instead a Markov kernel to , with its measurable structures specified. An optimization algorithm is an iteration of such steps, including its query, acceptance, and stopping rules; the name alone does not assert successful minimization.
Code inspection and representation-dependent costs are explicit auxiliary data when they affect the algorithm. A randomized program may equivalently use a deterministic step with its seed in . The oracle is the interface answering objective or derivative queries; a black-box algorithm knows the objective through those answers rather than through a closed formula.
Theorem 6.5
Universal global-jet factorization
Every optimizer of Definition 6.4 on smooth objectives has a unique optimizer on global holonomic jets,
For partial maps its domain is transported by . The same statement holds for stochastic steps, with the measurable structure on holonomic sections transported from . Stopping behavior and the entire iterated trajectory are preserved from matched initial runtime states and corresponding objective and auxiliary-input streams.
Proof. The inverse identities in Definition 6.2 give the displayed factorization. Every in the holonomic space equals , so any factorization must have the displayed value and domain; this proves uniqueness. For stochastic steps use the same substitution in every measurable output event. The transported measurable structures make this substitution measurable. Induction over steps preserves deterministic trajectories; integral composition preserves the stochastic trajectory law and stopping events.
This is the precise sense in which every specified objective-based optimization algorithm is an instance of the global jet construction. It includes derivative-free methods because the objective’s values are its zeroth-order component. The theorem is a representation result: infinite derivatives can be available without each algorithm actually computing or using them. It also holds with global order- holonomic sections for any ; increasing the order enriches the accessible differential operations, rather than being necessary to recover a function already retained at order zero. Extending an optimizer to nonholonomic formal sections is a separate modeling choice and is not determined by this theorem.
Proposition 6.6
Compositional equivalence of optimizer interfaces
Fix auxiliary data. Let have runtime state spaces as objects and total steps as arrows, with composition . Replacing by gives a category isomorphic to through . Partial steps and stochastic steps have the corresponding partial-map and kernel compositions.
Proof. Identity steps return ; associativity follows from substitution. Since is shared by both composed steps, . Precomposition with is inverse to precomposition with . The same argument transports domains, and kernel composition uses the integral law already specified.
Composition permits search, differentiation, filtering, acceptance, resource checks, and release gates to use one typed objective representation while retaining their separate runtime states.
Proposition 6.7
Criterion for a pointwise-jet optimizer
Fix and a total germ-local rule , with a set. It factors through realizable order- pointwise jets, including , exactly when its value is constant on germs having the same -jet, for each fixed . For a partial rule its domain must also be a union of these equivalence classes.
Proof. A factorization is constant on the fibers of the jet map. Conversely, assign to each realizable jet the value of any representative germ. The premise makes this assignment independent of representative, giving the unique descended rule. The same reasoning applies to a partial rule precisely when definedness also descends.
Here germ-local means dependence only on agreement on a neighborhood; pointwise-jet determined means dependence only on the derivative data at the base point. These are distinct conditions. For example, choose a smooth function supported in with . The objectives and have the same germ at zero. Searching the candidates returns zero for and one for . This legitimate value-based search cannot factor through the current germ or current pointwise jet. It does factor through the global holonomic jet section. Likewise the flat-germ example in Definition 6.1 explains why the infinite pointwise jet can forget information even in a germ-local problem.
Orders, queries, and optimizer state. Gradient methods read first-order jets of at their query points. Newton methods read second-order jets. A line search selects a step length along a declared direction by querying candidate values; it can be expressed through zeroth-order evaluations of the global jet section. Finite differences use several such evaluations to approximate derivatives. Quasi-Newton methods update an approximate Hessian or inverse Hessian from successive gradient and displacement pairs, so their memory belongs to . Momentum and Adam combine first-order queries with retained memory. Population and randomized search methods retain several candidate points and query values at each. A trust-region method optimizes a local Taylor model under a bound with ; higher-order tensor methods use higher multilinear derivatives. Constraints and acceptance tests remain part of the update interface.
For a composite loss , algorithms that use the network’s representation, such as Gauss–Newton, retain and as auxiliary data, or use their jet sections jointly. Equality of the scalar loss alone need not identify this representation. Minibatch methods likewise specify the sampled objective family and the batch-selection state. Chart changes act on jets by the chain rule, and coordinate-invariant training still requires transport of the optimizer geometry as in Proposition 4.2.
For nonsmooth or discrete objectives, classical infinite derivative jets may not exist. The universal value interface remains: a function is its global zeroth-order section. Subgradients, generalized derivatives, or surrogate derivatives can be additional declared oracle data. Consequently every objective-based algorithm has a global value/jet interface in its appropriate regularity setting. This does not identify arbitrary black-box search with a rule using only one smooth germ’s derivative sequence. The distinction is also computational: a runtime can expose finite jet queries on demand, without storing an infinite array or supplying an oracle for uncomputable objectives. A subgradient of a convex function at is a vector with for every in its domain. Convexity means for on a domain containing each such line segment. Other generalized derivatives mean explicitly chosen set-valued replacements for classical derivatives; a surrogate derivative is an explicitly supplied replacement used by the algorithm, for example from a smooth approximation.
For , the parameter Hessian is
Newton-type updates use this second jet. A damped positive-definite approximation can supply in Theorem 5.1, where is the declared symmetric Hessian approximation and damping is chosen so that is positive definite. Gauss–Newton retains the first term of (11); a raw indefinite inverse Hessian need not produce descent. Higher Taylor models, trust-region methods, and tensor methods can use larger jets, but a jet by itself is not an optimization algorithm: objective, step selection, constraints, and acceptance must also be chosen.
For an actual finite parameter step , Taylor’s theorem gives
with when the st derivative is bounded by along the connecting segment. This quantifies the step-size limitation of a finite jet description.
The phrase “jet tangent kernel” does not specify a single universal object. For a precise version, take continuous inputs , vector outputs, and multiindices . Define jet observables . Assume the mixed derivatives needed below exist and commute. Stack the chosen sampled observables into and set
Definition 6.8
Order-$r$ input-jet tangent kernel
For a parameter preconditioner , define the block kernel
The observable coordinates in this definition are unnormalized derivatives; using Taylor-normalized coordinates rescales the corresponding blocks.
For independent of input, differentiating the ordinary output block kernel gives .
Proposition 6.9
Jet-observable dynamics
For training on with , the stacked input jets satisfy and .
This is appropriate for derivative matching and Sobolev objectives; Sobolev training is an established example [8]. A sampled Sobolev objective, for example, is , with nonnegative weights and prescribed derivative targets . The term means that function values and chosen derivatives both enter the loss. It is distinct from the parameter-jet tower in (10) and the NTH in (9). One may also construct weighted Gram matrices from higher parameter derivatives, but without an optimizer law such matrices do not determine training.
For discrete tokens there are no ordinary input derivatives in the token index; input jets can instead be taken in continuous embedding variables when that is the intended model. Parameter derivatives of logits remain meaningful wherever the parametrization is differentiable.
For momentum, optimizer memory is essential: matched parameters can have different velocities. For Adam [20], with gradient ,
Here , the learning rate is , and regularizes division; operations are componentwise, squares each component, and with zero initial moments , are bias corrections. The update depends on history through , not solely on a kernel at . Minibatches, schedules, weight decay, and architecture changes add further state or drift. If a diffusion model for training is specified, , Itô’s formula gives
Here is standard Brownian motion, the specified drift vector, the noise-amplitude matrix, and the matrix trace. The Hessian correction is absent from deterministic gradient flow. ReLU, the rectified linear unit , and other nonsmooth mechanisms require piecewise-smooth analysis, generalized derivatives, or a chosen surrogate; higher classical jets need not exist on switching boundaries. Quantized updates and sampling require discrete semantics. Derivative-free optimization fits the global zeroth-order interface; it need not depend on pointwise derivatives. Quantization means a declared map to a finite-precision value set, and sampling draws outcomes from a declared law.
As Definition 6.1 shows, the infinite jet at one point need not determine a smooth germ. Theorem 6.5 instead uses the global holonomic section and the full optimizer state. Analyticity restores local determination by a convergent Taylor series; it does not remove optimizer or environmental dependence. The justified differential statement is the chain-rule identity for differentiable parametrized models under specified update laws. A single fixed “jet kernel” does not determine all training, inference, deployment, and organizational dynamics.
For an actor , choose an observation map and an admissible action family . A capability change may refine , enlarge , or alter resource access. In a set-level deterministic model, states are observationally equivalent to when . In higher semantics, use the homotopy fibers of and retain witnesses when appropriate. This describes capability through accessible observations and actions, with no assumption that capability is literally a homotopy level.
Theorem 7.1
Observation-fiber obstruction to regulation
Let be a deterministic transition, restrict to the reachable observation image , and let be the desired safe region. Define . A controller using only regulates all states only if
Conversely, a supplied selection from these intersections defines a regulator. In constructive type theory, the converse requires an actual dependent selection, not merely an assertion that each intersection is nonempty.
In particular, if indistinguishable states need incompatible actions, no extra computation applied to the same instantaneous observation can solve the problem. Memory or new sensors can change the observation object. If a finite disturbance set requires a different unique action for each disturbance, any regulating observation must distinguish every disturbance: , or at least bits of distinguishable observation. This restricted theorem illustrates requisite variety; it is not a blanket entropy inequality for every control system. Conant and Ashby’s good-regulator result has its own optimality and modeling assumptions [6].
An observational trace is a finite sequence of typed interaction events with declared times, participants, and outputs. An employee trace comes from work performed in an employee role; a customer trace comes from use of a product or service in a customer role. One person can occupy both roles. Data provenance specifies origin, transformations, and the permissions attached to the data. These traces become inputs to a learner only through a specified collection and data-processing map. This observational use of “trace” differs from an executable operation sequence.
Let be existing data, a newly captured work trace, and a future prediction target, all modeled as finite random variables on a declared probability space. The Shannon conditional entropy, in bits, is
Log loss for a predicted distribution is ; the Bayes-optimal predictor is the true conditional distribution, minimizing expected loss. Conditional mutual information is . Conditional relative entropy compares distributions through , averaged over the conditioning data. For finite random variables under log loss, the Bayes-optimal expected loss is conditional entropy. Hence the gain from adding the trace is
This follows directly from the definition of conditional mutual information and nonnegativity of conditional relative entropy. It gives a precise condition under which internal traces have predictive value. It does not establish equal value to customer data, a monetary valuation, or an incentive to capture every trace: acquisition cost, provenance, representativeness, permissions, and objective choice remain independent variables.
For worker and customer trace families , , equal predictive value would be the additional equation for the declared target, baseline, and data budget. It does not follow from both being traces. For example, with constant, a fair bit, an independent fair bit, and , these two quantities are zero and one bit, respectively. Information value, collection coverage, and authority must therefore be specified independently.
Definition 7.2
Observation refinement and authority
For a common state domain , observation maps and satisfy when for some map . Thus ’s view can be reconstructed from ’s. This is a preorder: reflexivity uses the identity and transitivity uses composition. Mutual refinement defines equivalence of information; the preorder is not assumed antisymmetric. For a common intervention interface , let be actor ’s authorized actions. For a common policy space , let be its authorized policy changes. Policies specify controllers, objectives, observation access, and update criteria. Authority domination means inclusion of the corresponding action and policy-change sets in each state. Strict domination means at least one inclusion is proper.
The channels and rights concern a declared organizational interface, rather than every fact an actor knows or every action an actor can perform. Controller inspection and challenges to updates are explicit observation or governance commands when included in this interface.
Definition 7.3
Trace coverage
Choose a finite operational state abstraction , including the relevant firm, environment, and interaction histories. Let be a finite set of employee and customer role instances, with weights and . For each choose a nonconstant trace observable , with finite, and an observation map . Write for the firm authority role and for its finite observation channel. Take each observation codomain to be its reachable image. Define
The number is informational trace coverage for these observables and weights; is a recovery witness. High coverage means for a declared threshold .
Recoverability is equivalent to being constant on each fiber of : this condition defines on the reachable image, and a recovery map immediately implies it. Nonconstant trace observables exclude uninformative constant functions from the coverage count. Recoverability does not by itself say that a particular storage or collection mechanism was used. For noisy channels one can instead declare a probability law and errors , requiring . Choosing below the best constant-predictor error ensures that this approximate recovery is informative.
Definition 7.4
Panoptic architecture and panoptic hypothesis
Let contain the state domain, roles, observation maps, trace observables, weights, actions, and governance relations just defined. It is a panoptic architecture at threshold , written , when:
- ;
- for every , and, in every state, and ;
- for some and state , at least one of these two authority inclusions is proper.
For a declared collection of firms and a fixed assessment period, the panoptic hypothesis is the empirical assertion . The literal claim about every AI company uses that entire declared class for ; a case study can instead test a restricted collection.
This is an operational definition inspired by the asymmetry of observation and authority in the historical account of panopticism [10]. It specifies the mathematical meaning of the term in this note. The threshold, role weights, trace scope, recovery errors, and authorization evidence are part of the hypothesis and must be reported. Positive predictive value alone proves none of its three conditions. A model can have full trace coverage and fail the definition when all action and policy-change rights are equal. Conversely, authority concentration does not establish high trace coverage. Controllability means the ability to reach declared target states through allowed inputs; observational predictability is a separate property.
Proposition 7.5
Panoptic status is invariant under structured equivalence
In the finite model, bijective changes of state, observation, trace, action, and policy coordinates preserving roles, weights, maps, and authorization relations preserve .
Proof. A recovery map transports by composing with the observation bijection’s inverse and the trace bijection. Hence exactly the same role instances contribute to . Observation factorizations transport in the same way. Bijections preserving the authorization relations preserve inclusion and proper inclusion of the corresponding sets. All three conditions are therefore preserved, in both directions.
In higher contextual semantics a recovery witness has the internal type
Mere recoverability is its -truncation, while the full type retains maps and coherence witnesses. Numerical coverage above uses a finite, decidable abstraction; a general higher recovery proposition is not automatically a decidable Boolean. An observation-only equivalence which forgets permissions need not preserve panoptic status. The structured equivalence must retain the authority data as well.
Deployment also changes data: write , , and . This feedback is related to performative prediction [28]. It can amplify, attenuate, or redirect effects; its sign and stability must be estimated, not inferred from the existence of a loop.
A company contains and connects decision-makers with distinct objectives. The set-based open-game formalism [14] represents a component by a strategy set , play , coplay , and a best-response relation
An equilibrium in context is a strategy with . Sequential and parallel composition retain the environmental continuation , rather than asserting that each isolated optimum remains optimal after composition. The type specifies the outcomes relevant to this component; returns outcomes to its environment. A deployment gate can thus expose actions forward and evaluation consequences backward. The formalism represents best responses, not a guarantee that a neural decision-maker implements them or converges to an equilibrium.
Hedges’s research proposal [16] highlights the dependence created when apparently separate economic decision-makers use a shared upstream AI provider. A minimal probabilistic example makes the distinction precise. Suppose a finite latent provider regime affects two otherwise conditionally independent action channels. Then
under the stated conditional-independence assumptions. This mixture generally does not factor into the two marginal channels. For and , it gives , whereas the product of marginals is . This is a shared-cause example, not evidence of actual collusion. Common providers, jointly updated policies, and correlated failures should therefore be represented as shared state in an industry model; corporate separation alone does not prove independence of decision channels.
An equilibrium is a state with zero time derivative. It is asymptotically stable when all sufficiently close initial states remain close and converge to it. For the linear systems below this is equivalent to every eigenvalue having negative real part. A kernel mode is an eigenvector with ; its eigenvalue is . Consider a local two-mode approximation with learning residual and an environmental or organizational response :
One possible frozen-kernel interpretation is for a learning mode, while encode the surrounding response and timescales.
Proposition 8.1
Two-mode closed-loop criterion
The equilibrium of (16) is asymptotically stable exactly when . If , it has a positive real eigenvalue.
Proof. The characteristic polynomial is . Its roots have negative real parts exactly when both coefficients after the leading term are positive. For negative determinant the roots are real and have opposite signs. At a zero eigenvalue prevents asymptotic stability.
Increasing an isolated learning rate, or making its loss decrease, does not by itself control and . Delays, switching releases, and nonlinear saturation require a richer model. A categorical wiring diagram specifies composition; it does not provide the missing stability estimates.
Resources can be made explicit through, for example,
with an available energy stock and a financial balance, together with material stocks and capacity constraints. Power is energy per unit time. A viability region includes resource floors as well as technical and institutional conditions. Maintaining that region is a control problem with disturbances; parameter-loss minimization proves only the narrower statement in Theorem 5.1.
Beer’s viable-system roles suggest a decomposition into operational model teams, coordination of shared infrastructure, allocation and audit, adaptation to the environment, and policy/identity [2]. This is a modeling correspondence, not evidence that an AI company is a biological organism or already a viable system in Beer’s technical sense.
Suppose compatible data objects admit a pushout in the chosen semantic category. This expresses a universal integration construction: for every test object ,
This defines the pushout as the gluing of the span through . It entails no improvement in statistical quality or profit. Permissions, shared identifiers, and interface compatibility are prerequisites for the construction to represent a real integration.
Let and be explicitly specified net objectives. Decompose their difference as
Here the terms are avoided duplication and complementary-data gains, while the terms are the respective additional costs. Under this accounting, integration is preferred for this objective precisely when the right side is positive. This is a conditional comparison, not a prediction: independent objectives, diseconomies, limited resources, failure concentration, and governance can reverse the sign. No date or inevitability of industry-wide consolidation follows from categorical universal properties or from (15).
The framework supplies a starting point with several measurable extensions. For a first case study, choose one workflow from employee or customer interaction to training, evaluation, release, and subsequent observations. Record its state variables, time delays, optimizer memory, permissions, resource costs, and rejection paths. Specify before analyzing updates. Identify which equivalences are intended to preserve predictions only, and which must also preserve optimizer geometry, costs, or control rights.
Estimate empirical tangent-kernel modes or Jacobian-vector products on a declared sample; materializing a full kernel is often unnecessary. Track kernel change, jet approximation error over actual step sizes, and sensitivity of behavior to release and data-selection decisions. Estimate the coupled response terms in (16); test whether broader observations resolve the obstruction in Theorem 7.1. These measurements separate information capture, effective intervention, and governance.
Four theoretical extensions follow naturally:
- Compositional certificates. Give module contracts and prove that their shared-resource and interface witnesses assemble into the descent datum of Proposition 3.3. Arbitrary local tests do not suffice.
- A moduli object of learners. Define equivalences preserving specified forward behavior, losses, optimizer geometry, and permissions. Compute stabilizers and ask which kernels and costs descend to its quotient stack.
- Multiobjective and stochastic governance. Replace a single scalar company objective by stakeholder-indexed objectives and a specified mechanism for selecting actions. Add Markov kernels, delays, and resource-aware policies.
- Jet closure with error control. Derive architecture- and optimizer-specific truncation estimates, rather than importing a smooth fully-connected NTH bound into an unrelated runtime.
Relating cocycles to homotopies requires a specified complex or derived moduli problem. A cochain complex has vector spaces and maps with . Cocycles are , coboundaries are , and cohomology is . A cochain homotopy between maps is a degree- family with . In an ordinary cochain complex, cocycles are not automatically null-homotopic: two cocycles represent the same cohomology class when their difference is a coboundary. For a smooth quotient by a Lie group, the infinitesimal action gives a two-term complex
in degrees . Here a Lie group is a smooth manifold with smooth group operations, is its tangent space at the identity, and is the derivative of the action there. The cokernel is the target vector space modulo the image. Its kernel records infinitesimal stabilizers and its cokernel records tangent directions modulo the action. A derived enhancement can introduce genuine obstruction and higher-homotopy information, but it must be constructed [29]; it is not supplied by ordinary Taylor coefficients.
Finally, a one-point model exists for many purely equational algebraic theories, including the theory of groups, but not for every consistent theory. The theory of nontrivial fields, with , is a counterexample. An external model gives relative consistency through a sound metatheory; it does not automatically give an executable or physically realizable model. The practical aim is a hierarchy of explicitly presented models and checkable claims connecting formal semantics to observed organizational behavior. A model interprets the operations and predicates of a theory so that its axioms hold. A purely equational theory specifies operations and equations between their terms; logical consistency means that a contradiction is not derivable. Soundness of a metatheory means its derivations preserve truth in the models under consideration. A derived deformation problem, proposed here as a further extension, retains homotopy-valued families of perturbations and their higher relations; its actual complexes and obstruction maps must be specified before any additional claim is proved.
- [1]W. Ross Ashby. An Introduction to Cybernetics. Chapman and Hall, London, 1956. Requisite variety: Chapter 11.
- [2]Stafford Beer. Brain of the Firm. John Wiley and Sons, second edition, 1981.
- [3]R. F. Blute, J. R. B. Cockett, and R. A. G. Seely. Cartesian differential categories. Theory and Applications of Categories, 22(23):622–672, 2009.
- [4]Matteo Capucci, Bruno Gavranović, Jules Hedges, and Eigil Fjeldgren Rischel. Towards foundations of categorical cybernetics. Electronic Proceedings in Theoretical Computer Science, 372:235–248, 2022.
- [5]Lénaïc Chizat, Edouard Oyallon, and Francis Bach. On lazy training in differentiable programming. Advances in Neural Information Processing Systems, volume 32, 2019.
- [6]Roger C. Conant and W. Ross Ashby. Every good regulator of a system must be a model of that system. International Journal of Systems Science, 1(2):89–97, 1970.
- [7]G. S. H. Cruttwell, Bruno Gavranović, Neil Ghani, Paul Wilson, and Fabio Zanasi. Categorical foundations of gradient-based learning. Programming Languages and Systems (ESOP 2022), LNCS 13240, pages 1–28. Springer, 2022.
- [8]Wojciech Marian Czarnecki, Simon Osindero, Max Jaderberg, Grzegorz Swirszcz, and Razvan Pascanu. Sobolev training for neural networks. Advances in Neural Information Processing Systems, volume 30, 2017.
- [9]Brendan Fong, David I. Spivak, and Rémy Tuyéras. Backprop as functor: A compositional perspective on supervised learning. 34th Annual ACM/IEEE Symposium on Logic in Computer Science, pages 1–13, 2019.
- [10]Michel Foucault. Discipline and Punish: The Birth of the Prison. Vintage Books, 1995. Translated by Alan Sheridan. “Panopticism,” pp. 195–228.
- [11]Tobias Fritz. A synthetic approach to Markov kernels, conditional independence and theorems on sufficient statistics. Advances in Mathematics, 370:107239, 2020.
- [12]Bruno Gavranović. Fundamental Components of Deep Learning: A Category-Theoretic Approach. PhD thesis, University of Strathclyde, 2024.
- [13]Bruno Gavranović, Paul Lessard, Andrew Joseph Dudzik, Tamara von Glehn, João Guilherme Madeira Araújo, and Petar Veličković. Position: Categorical deep learning is an algebraic theory of all architectures. Proceedings of the 41st International Conference on Machine Learning, PMLR 235, pages 15209–15241, 2024.
- [14]Neil Ghani, Jules Hedges, Viktor Winschel, and Philipp Zahn. Compositional game theory. Proceedings of the 33rd Annual ACM/IEEE Symposium on Logic in Computer Science, pages 472–481, 2018.
- [15]Alexander Grothendieck. Pursuing stacks. Circulated manuscript, 1983. Revised transcription edited by Mateo Carmona with Ulrik Buchholtz, 2021.
- [16]Jules Hedges. AI safety meets value chain integrity. Cybercat Institute, December 11, 2023. Research proposal.
- [17]Martin Hofmann and Thomas Streicher. The groupoid interpretation of type theory. In Twenty-Five Years of Constructive Type Theory, Oxford Logic Guides 36, pages 83–111. Oxford University Press, 1998.
- [18]Jiaoyang Huang and Horng-Tzer Yau. Dynamics of deep neural networks and neural tangent hierarchy. Proceedings of the 37th International Conference on Machine Learning, PMLR 119, pages 4542–4551, 2020.
- [19]Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks. Advances in Neural Information Processing Systems, volume 31, pages 8571–8580, 2018.
- [20]Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. International Conference on Learning Representations, 2015.
- [21]Ivan Kolář, Jan Slovák, and Peter W. Michor. Natural Operations in Differential Geometry. Springer, 1993.
- [22]Saunders Mac Lane and Ieke Moerdijk. Sheaves in Geometry and Logic: A First Introduction to Topos Theory. Universitext. Springer, 1992. Corrected reprint 1994. Internal language: Chapter VI.
- [23]Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington. Wide neural networks of any depth evolve as linear models under gradient descent. Advances in Neural Information Processing Systems, volume 32, 2019.
- [24]Peter LeFanu Lumsdaine. Weak ω-categories from intensional type theory. Logical Methods in Computer Science, 6(3:24), 2010.
- [25]Jacob Lurie. Higher Topos Theory. Annals of Mathematics Studies 170. Princeton University Press, 2009.
- [26]Per Martin-Löf. Intuitionistic Type Theory. Bibliopolis, Naples, 1984. Notes by Giovanni Sambin.
- [27]David Jaz Myers. Categorical systems theory. Book draft, last updated September 3, 2023.
- [28]Juan C. Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, and Moritz Hardt. Performative prediction. Proceedings of the 37th International Conference on Machine Learning, PMLR 119, pages 7599–7609, 2020.
- [29]J. P. Pridham. Unifying derived deformation theories. Advances in Mathematics, 224(3):772–826, 2010.
- [30]D. J. Saunders. The Geometry of Jet Bundles. Cambridge University Press, 1989.
- [31]David I. Spivak. Functorial data migration. Information and Computation, 217:31–51, 2012.
- [32]The Univalent Foundations Program. Homotopy Type Theory: Univalent Foundations of Mathematics. Institute for Advanced Study, 2013. See especially Chapters 2, 3, and 9.
- [33]Benno van den Berg and Richard Garner. Types are weak ω-groupoids. Proceedings of the London Mathematical Society, 102(2):370–394, 2011.
- [34]Vladimir Voevodsky. A very short note on homotopy λ-calculus. Unpublished note, 2006. Revised version dated 2009 in the author’s archive.
- [35]Norbert Wiener. Cybernetics: Or Control and Communication in the Animal and the Machine. MIT Press, second edition, 1961. First edition 1948.
- [36]Greg Yang. Tensor programs II: Neural tangent kernel for any architecture. Preprint, 2020.