Distilled models.
AbstractPhila PRO
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
posted an update about 2 hours ago
GPT, Gemini, Claude, and I have identified a multitude of direct utilities for Beatrix useful for diffusion conditioning in very powerful geometric formats. We have also identified multiple weaknesses to compensate for, multiple strengths to augment, the cause of the final layer's weak erank output state, and an emergent mathematical property of calculation in this format. The final stage directional magnitude is overwhelming and becoming amplitude.
There is a full article brewing for this information, including a massive set of information already learned from Beatrix V3 that could not be extracted from the 2s variant.
As the model trains, the amplitude begins to strengthen over and over. The weak tokenization processing from splat attention, forms the internal state of the model towards a bloating fashion. This is due to the articulation applied by the structure of the aleph addressing.
This creates massive erank geometry naturally, exhausting the space, producing comprehensively complex geometric structures. This internal structure here is weakly bound to the internal bytes, causing recon to weaken over time >2048, producing the output tokenization to be weaker at higher token lengths. Training improves this but is not known to solve it.
Along this chain the final layer has formed a sort of unexpected behavior, an amplitude behavior. I've met amplitude responses before in multiple models, and even attempted to curated magnitude through flow matching to some success, however amplitude in that nature is costly and adds additional overhead to the train so I'll need to come up with something more careful, and potentially something more clever than just attaching a composite or an energy dampener.
Attention will be solved by introducing various MHA layers throughout, ensuring the recon through the depth of the model survives. With that we'll want to ensure large erank composites form as well, allowing those humongous geometric structures to form and contribute. updated a model about 6 hours ago
AbstractPhil/alephllm-mini-beatrix-training repliedto their post 1 day ago
My apologies for the incorrect format for the AMOE arms from the experimental branch. They have been saving as torch objects. They are now correctly saving as safetensors format. My apologies for the inconvenience this may cause for you use. I will be modifying the codespaces to use the correct safetensors formats.
After the first 20.9b tokens trained, the real experiments begins. Beatrix V3's first prepped-state modular command structure has been attached for dynamic training. These arms will exist as appendages for Beatrix - trained alongside with the trunk until the end of the run.
These exist for experimental extraction, analysis, distillation experiments, memory experiments, mathematics experiments, and more. Each arm will be built along the chain for specific test cases. Expectation for each is already lined up and the outcomes are tested for, but the model still may face instability and must be monitored.
As her first arm learns tinystories, she builds direct composite semantic structure throughout this system. Think of it like, the first higher-functioning cognition attachment.
She's still very naïve and structurally unaware, so attaching new limbs is essentially extending a structure that is not yet finished forming. Nothing but fragments of issued information from an unknown source.
In this case, this structure has been tested hundreds of times to ensure she will not simply collapse during training by having this attached.