Beatrix — AlephLLM Chat
Three byte-level AlephLLM crafts and their arms
Seems we've run into engineering problems. The time degraded from 4 seconds per step to 9 seconds per step, compounding increase per independent arm attached.
I need to halt and run analysis. As it stands, we'll run until the end of Friday, then the engineering goes up on blocks.
I'm going to squeeze every bit of every byte from this engineering. Formulas up for scrutiny, system optimizations up for scrutiny, kernel optimizations, libraries, everything. Running until October 21st isn't reasonable. The models must be prepared correctly and reasonably to the mathematics, the results must be calculated correctly and the model trained to parity without faulty modifications during training that can impact the results.
Here's an interactive viewer for the internals of Mini-Beatrix-2.5s
I'll enhance it for 3 when it's ready.
As a comparison to global erank, we're looking at a structure of 400+ for around half of Beatrix V3 so far, so roughly 16+ blocks of erank >400, substantially stronger than the original two models for geometric attribution. The final block has a collapsing problem currently, but I believe others have the answer with autoregression models through a finalized projection smoothing layer concept. I haven't employed it yet though.
The fractal instability hits pretty early. You need rounding structures early otherwise the gradients explode at one point or another. The predominant problem was loss explosions. It happened because of ill-formed eigens in the intentional step structure I was experimenting with. 5 step cantor essentially ensured the cantor fractals deviate to a certain degree, and depth itself was meant to raise the steps of fractals to new states and interpolate the fractals.
If you use any of this, make sure you either pass it into AI for optimization - as it's likely terrible due to being my earlier models (I came from game development, optimization is very different). AI will be able to improve the speed and accuracy of the formulas.
https://github.com/AbstractEyes/geofractal/blob/main/src/geofractal/model/core/vit_beatrix.py
https://github.com/AbstractEyes/geofractal/blob/main/src/geofractal/model/positional/cantor.py
https://github.com/AbstractEyes/geofractal/blob/main/src/geofractal/model/core/geo_fractal_david.py
One of the problems was similarity. Almost everything was self similar, which in theory should have helped differentiate. However in practice, the structure found it's own similarity attractor basins that cause cascade corruption down the chain. The only solidity was to introduce eigen comparators through decomposition learning, which is a little different than autoregression. With this, the decomposition required more accuracy otherwise the system would always default to 1 of the first 3 steps - resulting in rigid or slightly less rigid articulations.
I measured fp64 being required for stable 4 step, and fp roughly 92 to be in a safe zone for stable Mandels at step 5. Julia requires something substantially larger than mandels. Fp64 is ENOUGH for rotary offset in standard positional systems, however fp128 is required for something akin to cantor fractal positional systems of differentiation.
It happens due to the eigenvalues themselves often malforming, and the subsystem silently rounds them. Using FULL SVD is a compositional fix for comparison, with that introduces a huge overhead as well.
Fractals themselves turned out to be more compositionally useful, not as additive elements, but as miniature rounding structures. The splat there was built under the concept of eigen substitution, meant to composite a series of tiny opinions from tons of subsystem residuals together into a composite "blackboard", forming a more robust and structural aligned INK BLOT splat, similar conceptually to viewing a random inkblot. This eventually composites into a utility of structural awareness, and it really doesn't take very long.
Essentially, that structure is geometric in nature, but it's not using Eigenvalues directly. It CAN use them, it should be capable of using any structural bounds with attributable contributions.
Splat functions viably at bf16, is a bit slower than MHA, but houses geometry more cleanly than MHA (sometimes by a huge margin) when trained with MUON instead of adam, adamw, or another multitude of optimizers I ran. I have attempted custom optimizers to encourage this behavior further, but the results showed MUON is just better.
Give it a shot in something simple, it'll train fast enough.
Pretty much anything in here is useful.
https://huggingface.co/collections/AbstractPhil/geolip-research-concepts
Eigens and causal chains have correlations but not causation without additional contributions to the assessments, the SVAE shows this to be a guarantee in many shapes, and in many others impossible.
The accuracy between the two requires a smoothing system, alpha differentiation through projected MHA-esque alpha attention to patchworks in order to fill the gaps. They don't directly line up quickly though, it looks more soupy when it's done.
They coalesce, but the extractions aren't consistent enough to directly use without a series of wrappers and structural alignment systems. Cantor Aleph and Omegas are essentially this structural system, but they are unstable. Cantor fractals remain unstable until around fp128 for Mandelbrot without redefining the underlying methods the mathematics linalg system uses. I did some headway on this, but I ran into a glacier that I would have needed to sink months into to make headway so I built a system to replace the slower linalg systems and the system lost much of it's cantor fractal capacity in favor of reproducibility and consistency.
The prototype forged from a 52,000 battery sweeps to find the most consistent recon convergence over time, heavily scrutinized and analyzed for over a month to build into something useful.
https://huggingface.co/AbstractPhil/geolip-SVAE
The current best case of the eigens research conclusions. Everything SVAE built to the attention prototype, everything constellation built to the processing and lookups for the model, everything distillation built to the banks for capturing and yielding, everything structural built from the knowledge and wisdom of other researchers. Well, not everything structural - it needed a lot of geometric formation and structural cohesion to make it work.
https://huggingface.co/AbstractPhil/aleph-splat-0/blob/main/splat_attention.py
Attempts to speed the SVD up were somewhat fruitful, somewhat not. They are good for inference, but I never programmed the gradient backprops for it.
https://huggingface.co/AbstractPhil/svd-triton
The mobius lens being a faster form wasn't strong enough as an activation system. I needed an architecture around it, not just an activation.
https://huggingface.co/AbstractPhil/mobiusnet-distillations/blob/main/make_chart_1.py