AbstractPhila PRO
AI & ML interests
Recent Activity
Organizations
CaptionBert unlike bertenstein is surviving direct scrutiny through the formal re-examination so far. The student CaptionBert is valid to the claims.
Next we'll be fitting a semi-larger variant of CaptionBert specifically to the larger dataset and the whitening procrustes alignment system, specifically aligned to a few prototypes of champion from the earlier sets.
I'll set up a colab notebook at first to see how long it takes, and then progress to a more utilizable series.
Bertenstein was a good prototype for alignment testing, however invalid the Bertenstein pairwise comparators were. The alignment structure did yield some serious respective and utilizable contextual semantics and relational operational capacity down the chain.
Bertenstein has been archived and labeled what it is, an early prototype that led to Procrustes whitening analysis, which led to a considerably more robust line of distillation elements.
There are still more elements to explore in the multimodal structure, an entire world of science there, and yet I don't have the patience for it today. Distillation today, multimodal behavioral entropic shared space on another tomorrow.
I'm making sure to capture all the excess features before I cull the pod. There are many clip-vit-b feature sets that didn't make it to the features repo, which will push up the repo's feature counts a considerable amount.
These will include all the missing vitb feature sets that weren't pushed correctly or were omitted by oversight on my part. I'm remedying them for a direct continuity for future usage.
The semi-successful run series on the VIT-B lineup is live and full of useful baseline distillation information for feature + InfoNCE distillation processing as well as direct feature distillation processing, direct InfoNCE distillation, and multiple other tested methods. https://huggingface.co/blog/AbstractPhil/geometric-memory-ft4
This article showcases the baseline utilization and benchmarks of the earlier experiment line's objective and loss structures tested on 12m features for the vit-b baseline. Not the strongest showcase, but the strongest of the champions did show some serious promise.
Next setup will be a directly aligned set based on the loss and objectives decided by the champions in the first runs, for the second run they operate in direct conjunction with the bert-8192 and captionbert-8192 distillation format directly on clip-vit-l features - this time we're including DINOv3 into the mix for it's high potency.
I'm currently extracting 4 clip-vit-l variants for the CC12m features and will be running the next series on the L size, which will give considerably more active and useful features overall within a smaller package.
The captionbert-8192 has a more unique and difficult to tune for pixel processing parity, but I will spend a few days making sure the smaller prototypes fit before I run the large experiments in order to build towards the larger objectives.
Primarily I need to ensure the memory bank aligns correctly and the constellation conforms to the anchors correctly, as this process was not micro managed enough for this run. The results are nonetheless useful and potent.
The process continues until we cover the entire constellation series.
Geometric Memory FT4 — Distill Against a Consensus, Ship a Rotation
20 trained tinyvits dropping in roughly 8 hours or so with recorded data and an article.
Sorry my mistake, 42 trained vits.
The results are rolling in. The series is coalescing into the necessary implications per structure aligned with the 10m cc12m extracted dataset.
This will help determine the best and fastest utilizable series from the geometric ablation and objective construction systems historically, and with that the organized documented results will be concatenated and organized by Claude Fable to the necessary potentials for each system.
With this, each potential arm for the AMOE-LORA system will be robustly tested for their distillation principles. Faster are ideal for rapid LORA convergence and slower are ideal for anchoring differentiation convergence, while moderate with a bit of MSE overfit are good for generation to an extent, while moderately low generalizable states are more ideal for generalization preservation as gated logical dichotomy structures for the moderate decisions.
It's a bit more complex than that, a lot more complex, but the results are rolling out and will be in the next article based on geometric memory.
CommonCaptions12m clip-vit-laion-vit array is almost ready with the first 8 arms of inference for testing and a new battery of analysis to run. This will continue likely until the middle of August, but we'll see if I can complete it sooner by throwing some money at it.
Time to expand the arms for the vit into the full sail. We'll be hitting every major vit multi-teacher approach, including the memory anchoring finetune structures as well.
The memory bank systems have been shown to refine trained models within a degree of accuracy. The genetic experiments, the structural berts, and the vit collectives all showcased the possibility of this system's capacity to expand already pretrained systems by attaching expansions to those.
https://huggingface.co/AbstractPhil/geolip-bert-8192
https://huggingface.co/AbstractPhil/geolip-clip-vit-large-patch14-ctx576
https://huggingface.co/AbstractPhil/geolip-clip-vit-large-patch14-ctx576-seq77
https://huggingface.co/AbstractPhil/geolip-clip-vit-bigG-patch14-ctx576-seq77
https://huggingface.co/AbstractPhil/geolip-bertenstein
https://huggingface.co/AbstractPhil/geolip-vit-large-x3 i think?
https://huggingface.co/AbstractPhil/geolip-vit-x34 didn't work, too many vits
https://huggingface.co/AbstractPhil/geolip-captionbert-8192
Each of these are a testament to the utility of this concept.
One of the prototypes will include a multimodal memory bank with directly gated and interconnected shared memory gates speaking another model's language, rather than just embeddings for a singular model. This gate will take in one or multiple types of model inputs and process those inputs into an entirely different model series' responses in the AMOE format.
I will also be experimenting with the AMOE-LORA fused with memory bank processing directly rather than just gate. The constellation was baked from the anchored memory bank originally but it did not meet the same sort of embedding accuracy. However, the constellation results built the AMOE eventually. First things first though, have to step back and hit all the angles with all the necessary tests for robustness.
By stepping back to the earlier memory bank and fusing it with the alephs, the upcoming experiments will provide some solid strong-ended tests. With that the rapid training of captionbert will hopefully be applied to this tiny vit. If surge activates, the process may be strong enough from the memory bank to provide the necessary distillation learning speed required to train the full collective with minimal hardware.
I have many many models to train to create the full Beatrix V3 prototype, however the list is expanding nicely in order to provide a full multimodal type agnostic behavior within a reasonable MOE structure.