Last updated: 2026-10-03

P
Postgraduate research

K-Blade Geometric Algebra as a Structured Inductive Bias, Not a Replacement for Scalar Tensors

Authors: Pat Parslow (Draft)

Abstract

We propose constraining selected large language model (LLM) parameters and activations to structured, Clifford-algebra-valued representations — multivectors built from scalar, vector, bivector, and higher-grade components — implemented throughout by ordinary scalar coefficient tensors. Scalar tensors are not eliminated by this move; they remain the physical substrate. What changes is the multiplication rule imposed on them: the geometric, inner, and outer products of geometric algebra (GA), rather than an unconstrained linear map. We show how common neural primitives — linear maps, attention, positional encodings, and nonlinearities — can be expressed and implemented this way, define an attention similarity score that behaves correctly under GA's reversion operation, and treat grade truncation and closure under multiplication as a central design choice rather than an implementation detail. The strongest version of this proposal is narrower than a general replacement for scalar tensors: a Clifford-structured parameterisation is a plausible inductive bias specifically where a task involves compositional transformations or oriented subspace relations, and the right test of that claim is whether such a model generalises better per parameter than an unconstrained scalar linear map of equal size, not whether geometric algebra can express the operation at all. We present the mathematical mappings, a practical implementation, a microbenchmark establishing kernel-level feasibility (not yet model-level advantage), and a four-phase experimental programme designed to test the narrower claim directly, including a shuffled-algebra control to separate the effect of parameter sharing from the effect of geometric structure.

1. Introduction

Modern LLMs are built from large dense tensors containing scalar parameters and activations. While highly successful, scalar tensor parameterisations are agnostic to geometric structure: they do not capture rotations, subspace structure, or oriented subspace interactions natively. Geometric algebra (GA) provides a compact algebraic framework that unifies scalars, vectors, bivectors, and higher k-blades into a single graded multivector algebra equipped with the geometric product, outer (wedge) product, and inner product[1][2]. We explore replacing scalar tensors with k-blade based multivector tensors and adapting neural primitives to operate in this algebra.

2. Background FoundationalKnowledge that endures for decades — core principles

2.1 Geometric Algebra, blades, and multivectors — three distinct things

Geometric algebra extends vector spaces by forming multivectors: linear combinations of basis blades. These three terms are not interchangeable, and the distinction matters for everything that follows:cf. vectors, quaternions and blades

  • A k-blade is a decomposable homogeneous grade-k element, \(B = v_1 \wedge v_2 \wedge \cdots \wedge v_k\) for some vectors \(v_1, \dots, v_k\) — it represents a single oriented k-dimensional subspace.
  • A general grade-k element need not be a blade once \(n \geq 4\): a sum of two bivectors spanning different planes is a grade-2 element but is not, in general, itself decomposable into a single wedge product of two vectors.
  • A multivector is a sum of elements across several grades at once, e.g. \(M = M_0 + M_1 + M_2\).

Three further facts follow directly and constrain every design choice in Sections 3–4: the sum of two k-blades need not itself be a blade; the geometric product of two blades generally spreads across several grades, not one; and neither an attention-weighted sum of blade-valued values nor a coefficient-wise nonlinearity will generally preserve blade decomposability. A representation that must survive attention and a nonlinearity and still be called "k-blade-valued" in the strict sense is a much stronger claim than a representation built from grade-restricted multivectors. Key operations:

  • Geometric product: for vectors a,b

    \[ ab = a \cdot b + a \wedge b \]

    where \(a \cdot b\) is the symmetric inner product (scalar) and \(a \wedge b\) is the antisymmetric outer (wedge) product (a bivector).
  • Wedge (outer) product \(a \wedge b\): produces a higher-grade blade representing the subspace spanned by a and b.
  • Reversion \(\widetilde{M}\) reverses the order of the vector factors in each basis blade, which multiplies a grade-k basis blade by \((-1)^{k(k-1)/2}\): grades 0 and 1 are unchanged, grade 2 negates, grade 3 negates, grade 4 is unchanged, and so on. Reversion and grade projection together allow decomposition into scalar/vector/bivector parts, and reversion specifically is what makes \(\langle M\widetilde{M}\rangle_0\) behave as a genuine squared norm (Section 4.2).

2.2 Multivector expansion, and why scalar tensors are not eliminated

A general multivector M in GA(n) expands over basis blades {e_I} indexed by bitmasks I:

\[ M = \sum_{I \subseteq \{1,\dots,n\}} m_I \, e_I \]

Every coefficient \(m_I\) is an ordinary scalar, so the physical representation of M is still a tensor of scalars, now carrying an extra component axis of size up to \(2^n\). Nothing proposed in this paper removes scalar tensors from the implementation; what it proposes is constraining how those scalars are combined — through the geometric product's fixed multiplication table rather than an unconstrained linear map — and giving the resulting coefficient groups a geometric reading (direction, oriented plane, and so on) that an arbitrary hidden dimension does not have. Look at Section 6.1's own data layout and the choice is already made: one coefficient stored per retained basis blade, a grade-restricted multivector representation rather than literal k-blades in the strict decomposable sense of Section 2.1 — and the rest of this paper is written accordingly.

2.3 Tensors in Neural Networks

Neural networks use tensors of rank 1..N to store activations and parameters. Linear layers, convolutions, and attention are all implemented by scalar-linear algebra operations on tensors. Common optimisations exploit low-rank structure, parameter factorisation, and equivariant parameterisations when symmetry is present.

3. Conceptual Mapping: Unconstrained Scalars → Grade-Restricted Multivectors FoundationalKnowledge that endures for decades — core principles

3.1 Representational shift, and which of three architectures this actually is

Instead of storing a scalar per tensor element, store a small multivector per element. At minimum, each original scalar can be interpreted as the grade-0 (scalar) component of a multivector, while richer encodings place information across grades: vector components (grade-1) for directional features, bivector components (grade-2) for oriented plane interactions, etc. "Store a multivector" can mean three different architectures, and a design has to commit to one:

  • Blade-constrained: every representation must remain a decomposable blade. Strongest geometric interpretation; requires a special parameterisation or an explicit projection back onto the blade manifold after every operation that could break decomposability (which, per Section 2.1, is most of them).
  • Grade-restricted multivector: representations may be any element of a fixed set of grades, e.g. \(M = M_0 + M_1 + M_2\), without requiring any single grade to be a blade. Easier to train; this is what the rest of this paper means whenever it refers to a k-blade representation.
  • Full multivector: all \(2^n\) components are retained. Algebraically closed under the geometric product (Section 4.1), but component count grows exponentially in n.

This paper adopts the grade-restricted option throughout. The remaining sections should be read with "k-blade representation" standing for that choice, not for literal, strictly decomposable blades.

3.2 Structured embeddings

Word and token embeddings become multivector-valued embeddings: \(E(t) \in \mathrm{GA}(n)\) where GA(n) is built over an n-dimensional base vector space. Embeddings now encode magnitude (scalar), direction (vector part), and oriented subspace relationships (higher-grade parts). This can express richer relational priors (e.g., word analogies as oriented subspace rotations).a hypothesis. the runs below only time kernels

4. Neural Primitives in Geometric Algebra FoundationalKnowledge that endures for decades — core principles

4.1 Geometric product and linear maps

The geometric product between two multivectors generalises linear maps. For multivector parameters A,B and multivector input x, one canonical parameterisation is

\[ y = \sum_i A_i \, x \, B_i + b \]

where juxtaposition denotes the geometric product. Projecting grade components yields outputs in desired grades. Expanding in the scalar parameter space shows this is a structured block-linear map; the GA parameterisation can be much lower-dimensional when many blade coefficients are zero (k-blade sparsity). Whether it actually is lower-dimensional at equal expressive power is a separate, empirical question, taken up directly in Section 5.1.

Grade truncation, taken up only briefly in Section 8 below, is a defining property of this map, not an implementation detail. If only grades up to \(k_{\max}\) are retained, the representation is generally not closed under the geometric product: multiplying two bivectors, for instance, can produce scalar, bivector, and grade-4 parts at once, and if grade 4 is excluded the architecture must apply an explicit projection after every product,closure is a design choice, not a detail to fix later

\[ M_{\text{out}} = \Pi_{\mathcal{G}}\!\left(M_1 M_2\right), \]

where \(\Pi_{\mathcal{G}}\) discards every grade outside the retained set \(\mathcal{G}\). That is a well-defined neural operation, and there is nothing wrong with using it — but it carries three consequences that an architecture built this way inherits whether or not they are stated up front. Information is discarded on every product, not just once at the output. Once a projection sits between two products, associativity can be lost, since \(\Pi_{\mathcal G}(\Pi_{\mathcal G}(M_1 M_2) M_3)\) need not equal \(\Pi_{\mathcal G}(M_1 \Pi_{\mathcal G}(M_2 M_3))\): repeated products can behave differently depending on the order they are composed in. And the truncated structure that results is no longer the original Clifford algebra; it is a new algebraic object whose properties have to be established on their own terms, not inherited from GA's.

4.2 Attention

Following the query/key/value attention formulation[5], let queries, keys, and values be multivectors \(q, k, v \in \mathrm{GA}(n)\). The similarity score needs to use reversion to behave correctly as a general-multivector inner product:for plain vectors this scalar part is the dot product

\[ s(q,k) = \frac{\langle q\,\widetilde{k} \rangle_0}{\tau}, \qquad a_{ij} = \operatorname{softmax}_j\!\big(s(q_i, k_j)\big) \]

Without the reversion on \(k\), different grades can pick up a grade-dependent sign under the plain geometric product's scalar part, and the self-similarity \(s(q,q)\) is not guaranteed to behave like a norm. With it, \(\langle q\widetilde{q}\rangle_0 = \sum_I m_I^2\) (up to the fixed sign each basis blade squares to under the chosen metric signature), which is non-negative in the Euclidean case, exactly the property a self-similarity score should have. For grade-1-only queries and keys this reduces to the ordinary dot product, since reversion fixes vectors: \(\widetilde{k} = k\), recovering standard scaled dot-product attention as the special case where only the vector grade is populated. The scaling term also deserves more thought than copying \(\sqrt{d}\) from ordinary attention: the appropriate denominator may depend on the number of active blade components, per-grade variance, the algebra dimension n, the model width outside the algebra, the metric signature, and how the query/key operators are initialised. A practical default is to learn \(\tau\) rather than fix it, possibly per attention head or per grade.

The attended multivector is

\[ o_i = \sum_j a_{ij} \, v_j \]

4.3 Rotors and positional encodings

Rotations in GA use rotors \(R = \exp\!\left(-\tfrac{1}{2} B\right)\) for a bivector \(B\). Rotor action on a multivector \(M\) is

\[ M' = R \, M \, \widetilde{R} \]

where \(\widetilde{R}\) is the reverse of \(R\). This recovers rotary-style embeddings (RoPE)[4] as special cases and generalises them to higher-grade interactions.

4.4 Nonlinearities

Nonlinear activations for multivectors can be constructed by operating on invariants or per-grade components. Options:

  • Apply scalar nonlinearity to magnitudes: \(\sigma(\lVert \operatorname{grade}_1(M) \rVert)\) and rescale directions.
  • Apply elementwise scalar nonlinearity to each blade coefficient.
  • Use learnable grade-wise gates \(g_g\) per grade \(g\): \(\operatorname{grade}_g(M) \mapsto g_g \odot \sigma(\operatorname{grade}_g(M))\).

5. Theoretical advantages, and where they are conditional Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

5.1 Claimed advantages

  • Geometric expressivity: captures oriented subspace relationships naturally.
  • Parameter efficiency: conditional, not automatic. Replacing each scalar feature with L blade coefficients increases raw activation memory by roughly a factor of L; any net saving has to come from restricting the operator family, sharing coefficients across the \(A_i, B_i\) terms in Section 4.1's map, or genuine blade sparsity, not from geometric algebra as such. A fair test holds at least one of three budgets fixed when comparing against a scalar baseline — equal parameter count, equal FLOPs, or equal wall-clock training time — since a GA model outperforming a scalar model with the same nominal hidden width, but a different parameter count, does not by itself establish parameter efficiency.
  • Rotational and subspace equivariances: rotors provide mechanisms for equivariant transforms useful in language analogies and structured transformations.
  • Interpretability: grades map to geometric concepts (direction, plane) enabling richer probes — though this is itself a hypothesis to test, not a property guaranteed by construction (see Section 7, Phase 4).

5.2 Where this inductive bias is most likely to help

A full replacement of every tensor throughout an LLM, as the earlier framing of this paper suggested, is probably the highest-risk starting point: it bets the entire architecture on geometric structure being present in ordinary token embeddings, for which there is no particular guarantee. The idea is considerably more defensible as a targeted inductive bias applied where a task plausibly contains compositional transformations or oriented subspace relations:

  • Positional transformations, since rotor actions have a direct structural connection to rotations (Section 4.3).
  • Relation-aware attention heads, especially where direction, order, inversion, or composition matters.
  • Knowledge-graph relations, where geometric products could model compositional transformations between entities.
  • Small, deliberately constructed semantic subspaces, where bivectors represent interactions between a handful of learned directions.
  • Multimodal models, where actual spatial or rotational structure is present in at least one modality.
  • Agent-state and transformation models, where actions are naturally operators acting on a structured state.

Plain natural-language token embeddings are the weakest case for this bias: nothing guarantees that the latent factors language models discover have the rotational symmetry GA is built to exploit. The decisive question for any application is whether the geometry matches a real regularity already present in the data, not whether GA supplies an elegant way to describe the computation.

6. Practical implementation Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

6.1 Data layout and memory

Implement multivector tensors as arrays shaped (batch, seq_len, components) where components correspond to the chosen basis blades up to grade k_max. For base dimension n, the number of basis blades is 2^n; choose small n and restrict grades (e.g., up to grade 2 or 3) to cap memory.each extra dimension doubles the blade count

6.2 Efficient geometric product

The geometric product reduces to signed sums over basis blade products with a deterministic sparsity pattern. Implement via sparse kernels exploiting antisymmetry and grade structure. Two routes:

  • Precompute multiplication tables for the chosen basis and implement via batched sparse-dense matmuls.
  • Implement small dense microkernels (fixed small n) optimised on GPU via CUDA or Triton.

6.3 Initialisation and training

Initialise multivector parameters to match scalar baselines by placing scalar baseline weights on the grade-0 components and small random noise on higher-grade components. Train with standard optimisers; consider separate learning rates for higher-grade components.

6.4 Backward compatibility

Provide adapter layers to convert between scalar and multivector tensors, enabling staged adoption in existing transformer stacks (replace embeddings/linear layers first, then attention).

7. Experimental protocol Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The core testable hypothesis this protocol is built to check is narrower than a general case for GA: a Clifford-structured parameterisation provides a useful inductive bias specifically when a task contains compositional transformations or oriented subspace relations, enabling better generalisation per parameter than an unconstrained scalar linear map. That is falsifiable in a way "k-blade representations are powerful" is not, and the four phases below are ordered to test the smallest, cheapest version of the claim first.

Phase 1: Representation test

Before touching a language task, compare equal-dimensional representations on synthetic problems with known geometric structure: rotation composition, subspace intersection, orientation reversal, transformation analogy, and recovering a latent plane from noisy observations. The purpose is narrow — confirm the architecture can exploit the structure it was designed to express at all — before spending a training run finding out the hard way.

Phase 2: Controlled language component

Replace exactly one transformer component at a time against a fixed scalar baseline: multivector embeddings only; GA query/key similarity only; rotor positional encoding only; GA value transformations only. This isolates whether any gain comes from the representation, the similarity measure, or the structured transformation, rather than crediting a full-system swap for a win that one component produced.

Phase 3: Fair architecture comparisons

Compare against an ordinary dense model, a wider scalar model matched for parameter count, a low-rank model, complex-valued and quaternion alternatives, and a generic grouped linear layer — holding at least one of the three budgets from Section 5.1 fixed throughout. Include one further control specifically: a Clifford model with its multiplication-table signs randomly shuffled. If performance survives that shuffle, the benefit is most likely coming from parameter sharing and structured sparsity, not from the geometry itself, and the paper's claim should be revised accordingly rather than defended past the evidence.

Phase 4: Interpretability test

Do not assume grades are interpretable by construction. Test directly whether vector components consistently encode directional factors, whether bivectors correspond to identifiable relation planes, whether rotor parameters compose in semantically meaningful ways, and whether intervening on a specific grade produces a predictable behavioural change.

8. Limitations and pitfalls

  • Overhead: naive multivector expansion increases memory (exponential in n). Truncation and sparsity are essential.
  • Closure: truncating to grades up to \(k_{\max}\) breaks closure under the geometric product (Section 4.1) and can break associativity once projection is interleaved between products — a structural property of the chosen algebra, not a bug to patch later.
  • Optimisation: nonstandard parameter geometry may require tailored optimisers.
  • Hardware: high-performance GA kernels may need custom CUDA/Triton implementations.
  • Unclear universality gains: empirical validation required, and Section 7's shuffled-algebra control is specifically designed to catch a result that looks like a geometric win but is actually a parameter-sharing win.
  • Hypercomplex networks (quaternions, octonions) applied to RNNs/CNNs[3].
  • Clifford networks and geometric deep learning literature exploring equivariant representations, e.g. Clifford group equivariant neural networks[6].
  • Rotary embeddings[4] and complex-valued attention mechanisms.

10. Conclusion

Section 2.2 already settled what this architecture physically consists of: scalar coefficient tensors, the whole way through. The actual contribution sits one level up, in the multiplication rule imposed on those coefficients — a geometric product with a fixed, meaningful table, rather than an unconstrained linear map — and in the rotors, grades, and reversion-correct inner product (Section 4.2) available for building structured operators once that rule is in place. Framed that way, the claim is narrower than a general replacement for scalar tensors in LLMs, and considerably easier to test: a Clifford-structured parameterisation earns its place exactly where it generalises better per parameter than an unconstrained scalar linear map on a task with genuine compositional or oriented-subspace structure. Sections 5.2 and 7 are built to put that single comparison in front of the evidence, rather than take it for granted.

Acknowledgements

Draft prepared by Pat Parslow. Feedback welcome.

  • Algorithm, Not Metric — the empirically-tested counterpart to this paper's theoretical proposal: k-blade subspace similarity validated against a real 160-page corpus, including a failure mode (chaining) worth knowing before building on the representation proposed here.
  • Vectors, Quaternions, and Blades: Why Geometric Algebra Subsumes Both — works through, from first principles, why blades and rotors subsume vector and quaternion algebra, the same structural claim this paper leans on for rotary positional encodings (Section 4.3).
  • Understanding Large Language Models — a no-code account of the attention and embedding mechanisms this paper proposes re-expressing in geometric-algebra terms.

References

  1. D. Hestenes, New Foundations for Classical Mechanics, D. Reidel Publishing Company, 1986. https://doi.org/10.1007/978-94-009-4802-0
  2. L. Dorst, D. Fontijne, S. Mann, Geometric Algebra for Computer Science: An Object-Oriented Approach to Geometry, Morgan Kaufmann, 2007. https://geometricalgebra.org/
  3. T. Parcollet, M. Morchid, G. Linarès, "A survey of quaternion neural networks," Artificial Intelligence Review 53(4), pp. 2957–2982, 2020. https://doi.org/10.1007/s10462-019-09752-1
  4. J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, Y. Liu, "RoFormer: Enhanced Transformer with Rotary Position Embedding," arXiv:2104.09864, 2021. https://arxiv.org/abs/2104.09864
  5. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, "Attention Is All You Need," Advances in Neural Information Processing Systems 30 (NeurIPS 2017). https://arxiv.org/abs/1706.03762
  6. D. Ruhe, J. Brandstetter, P. Forré, "Clifford Group Equivariant Neural Networks," Advances in Neural Information Processing Systems 36 (NeurIPS 2023). https://arxiv.org/abs/2305.11141

Appendix: Minimal GA kernels (sketch)

Represent multivector as vector of blade coefficients indexed by mask 0..2^n-1 restricted to popcount(mask) <= k_max. Precompute blade multiplication table (sign, result_mask) and implement batched gather-add kernels on GPU. For fixed small n consider hand-optimised microkernels.

A. GPU microbenchmark (RTX 5070 Ti) Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates

To validate a simple batched geometric-product kernel, we implemented a PyTorch scatter-add based kernel and measured performance on an NVIDIA GeForce RTX 5070 Ti (CUDA 13.2 enabled PyTorch wheel).

Benchmark configuration

  • base dimension n = 3, grade truncation k_max = 3 (full GA for n=3)
  • basis size L = 8 (2^3 basis blades)
  • number of nonzero product pairs (precomputed) = 64
  • implementation: vectorised gather and scatter_add on CUDA (see ga_gpu_benchmark.py)

Results (averaged over 200 runs after warmup)

B,L,num_pairs,avg_ms,ops_per_sec

1,8,64,0.0466 ms, 21474.8 ops/sec

8,8,64,0.0528 ms, 18928.2 ops/sec

32,8,64,0.0433 ms, 23089.0 ops/sec

128,8,64,0.0612 ms, 16330.2 ops/sec

A.1 Triton microkernel benchmarks (broader sweep)

We extended the microbenchmark to several (n,k) settings and batch sizes. The CSV results are saved at /home/p/ga_triton_final_results.csv. Key aggregated statistics:

  • (n=3, k_max=3) atomic Triton ~0.24 ms, baseline ~0.26 ms
  • (n=4, k_max=3) atomic Triton ~0.28 ms, baseline ~0.30 ms
  • (n=4, k_max=4) atomic Triton ~0.28 ms, baseline ~0.33 ms
  • (n=5, k_max=3) atomic Triton ~0.36 ms, baseline ~0.40 ms

Representative per-case numbers (B, L, pairs, baseline_ms, opt_ms, atomic_ms, err_opt, err_atomic):


n,k,B,L,num_pairs,baseline_ms,opt_ms,atomic_ms,err_opt,err_atomic
3,3,1,8,64,0.2779,0.3398,0.2491,4.768372e-07,4.768372e-07
3,3,8,8,64,0.2617,0.3519,0.2469,4.768372e-07,4.768372e-07
3,3,32,8,64,0.2551,0.3516,0.2434,9.536743e-07,1.430511e-06
3,3,128,8,64,0.2606,0.3494,0.2383,9.536743e-07,9.536743e-07
4,3,1,15,211,0.3360,0.4554,0.2599,3.076522e+00,9.536743e-07
4,3,8,15,211,0.3001,0.4373,0.2774,3.292984e+00,9.536743e-07
4,3,32,15,211,0.2954,0.4308,0.2734,3.680134e+00,1.907349e-06
4,3,128,15,211,0.2939,0.4551,0.2894,6.226732e+00,2.384186e-06
4,4,1,16,256,0.3052,0.4816,0.2816,4.768372e-07,9.536743e-07
4,4,8,16,256,0.3480,0.4910,0.2744,1.907349e-06,1.907349e-06
4,4,32,16,256,0.3286,0.4711,0.2859,1.907349e-06,2.861023e-06
4,4,128,16,256,0.3057,0.4791,0.2817,3.337860e-06,1.907349e-06
5,3,1,26,556,0.3952,0.7072,0.3576,6.064522e+00,9.536743e-07
5,3,8,26,556,0.3853,0.6952,0.3627,1.006598e+01,2.384186e-06
5,3,32,26,556,0.4054,1.1341,0.3926,8.531096e+00,2.861023e-06
5,3,128,26,556,0.5594,0.7187,0.3663,1.247142e+01,3.337860e-06

A.2 Reading this benchmark correctly

This table establishes kernel-level feasibility, not model-level advantage, and both halves of that distinction matter. On the useful side: small geometric products execute efficiently, and the atomic Triton implementation stays close to baseline numerical error (around \(10^{-6}\)) across every configuration tested. On the cautionary side, the table also contains its own warning: the alternative "opt" implementation has large errors — err_opt running from roughly 3 up to over 12 — for every truncated case shown (\(n=4, k_{\max}=3\) and \(n=5, k_{\max}=3\)), exactly where grades are being discarded per Section 4.1's closure discussion, while the atomic implementation's error stays at the baseline's own numerical-precision floor throughout. Whatever implementation is used downstream of this appendix, it should be the atomic one, not the faster-looking but numerically unreliable "opt" path, on any configuration with truncation.

A second pattern in the data: the baseline, opt, and atomic timings are all broadly flat in the sub-millisecond range as batch size B runs from 1 to 128 (e.g. 0.2779 ms to 0.2606 ms for \(n=3,k=3\)). That flatness suggests kernel-launch overhead, not the geometric product itself, dominates these measurements — the batch dimension isn't yet doing enough work to show through. None of this table shows that a transformer layer built this way runs faster end to end, that training uses less memory, that the model needs fewer parameters at equal quality, that the resulting representation generalises better, or that it stays interpretable after training; demonstrating any of those is the job of the four-phase protocol in Section 7, not this appendix. Before any speed claim beyond "the kernel itself is cheap" is warranted, the next step is comparing larger, fused workloads directly against highly optimised dense matrix multiplication.

Final benchmark plots

Latency vs L
Figure A.1: Latency vs L (lower is better). Source: /home/p/ga_triton_final_results.csv
Speedup vs L
Figure A.2: Speedup vs L (higher is better). Source: /home/p/ga_triton_final_results.csv