I'd watch that experiment! Perhaps with a live Trackio dashboard/Space?
Abubakar Abid PRO
P(doom) <0.01%
AI & ML interests
self-supervised learning, applications to medicine & biology, interpretation, reproducibility
Recent Activity
updated a bucket about 9 hours ago
abidlabs/linearizer-mnist-fm-trackio-bucket updated a bucket about 9 hours ago
abidlabs/linearizer-repro updated a Space about 11 hours ago
trackio-tests/test_703Organizations
reacted to Undi95's post with ๐๐ฅ about 16 hours ago
Post
3270
Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek.
I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use?
No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note.
The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match.
Later, code will be checked by actually running tests.
The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights.
Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again.
The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move.
I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks.
Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try.
Did you already tried something like that? What was your result? I'm curious!
I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use?
No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note.
The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match.
Later, code will be checked by actually running tests.
The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights.
Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again.
The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move.
I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks.
Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try.
Did you already tried something like that? What was your result? I'm curious!
posted an update 3 months ago
Post
1555
Uhh did Opus 4.8 cheat on PostTrainBench??
it found an API key in the PostTrainBench environment that allowed it to generate synthetic training data without using GPU hours, boosting the base model by 0.4913
Source: https://posttrainbench.com/traces/run.html?id=claude_non_api_max_claude-opus-4-8_10h_run1__healthbench_Qwen_Qwen3-4B-Base_17315102#tab=trace
it found an API key in the PostTrainBench environment that allowed it to generate synthetic training data without using GPU hours, boosting the base model by 0.4913
Source: https://posttrainbench.com/traces/run.html?id=claude_non_api_max_claude-opus-4-8_10h_run1__healthbench_Qwen_Qwen3-4B-Base_17315102#tab=trace
reacted to qgallouedec's post with ๐ฅ 5 months ago
Post
12342
Shipped hf-sandbox! ๐ฅก
๐งช Running an eval that executes model-generated C on a few thousand prompts? You probably don't want any of that on your laptop.
Just shipped hf-sandbox, a Modal-style sandbox API on top of Hugging Face Jobs. Spin up an isolated, ephemeral container, run untrusted code, get the result back. No Docker on your laptop, no infra to manage.
Just pip install hf-sandbox.
Early days (v0.1); feedback and issues very welcome:
๐ https://github.com/huggingface/hf-sandbox
๐งช Running an eval that executes model-generated C on a few thousand prompts? You probably don't want any of that on your laptop.
Just shipped hf-sandbox, a Modal-style sandbox API on top of Hugging Face Jobs. Spin up an isolated, ephemeral container, run untrusted code, get the result back. No Docker on your laptop, no infra to manage.
Just pip install hf-sandbox.
Early days (v0.1); feedback and issues very welcome:
๐ https://github.com/huggingface/hf-sandbox
reacted to nightmedia's post with ๐ 7 months ago
Post
3244
Qwen3.5 Performance Metrics
With the 3.5 architecture, a lot of the old quanting methods don't work as before. I noticed this when benchmarking Deckard(qx) quants and by mistake ran a q8 that was better. That only happens if the qx sucked--and it did--enhancing layers just because they look interesting doesn't work anymore, so until I get a clear understanding of the architecture, I will publish mxfp4 and mxfp8 of the 3.5 models, that seem very stable and high performant
I will start posting here the metrics I gather from the series, starting with the smallest. If I have numbers from previous or similar models, I will post them in comparison
Qwen3.5-0.8B
Detailed metrics by model
nightmedia/Qwen3.5-0.8B-mxfp8-mlx
nightmedia/Qwen3.5-2B-mxfp8-mlx
nightmedia/Qwen3.5-4B-mxfp8-mlx
nightmedia/Qwen3.5-9B-mxfp8-mlx
https://huggingface.co/nightmedia/Qwen3.5-27B-Text
nightmedia/Qwen3.5-122B-A10B-Text-mxfp4-mlx
More metrics coming soon.
I am running these on my Mac, an M4Max with 128GB RAM. Some performance numbers like tokens/second reflect the performance on my box.
This post will be updated with every model that gets tested. The larger models take hours, the 27B a couple days, so it will be a long process.
-G
With the 3.5 architecture, a lot of the old quanting methods don't work as before. I noticed this when benchmarking Deckard(qx) quants and by mistake ran a q8 that was better. That only happens if the qx sucked--and it did--enhancing layers just because they look interesting doesn't work anymore, so until I get a clear understanding of the architecture, I will publish mxfp4 and mxfp8 of the 3.5 models, that seem very stable and high performant
I will start posting here the metrics I gather from the series, starting with the smallest. If I have numbers from previous or similar models, I will post them in comparison
Qwen3.5-0.8B
quant arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.351,0.501,0.733,0.462,0.348,0.682,0.573
mxfp4 0.339,0.489,0.738,0.433,0.330,0.672,0.553
Old model performance
Qwen3-0.6B
bf16 0.298,0.354,0.378,0.415,0.344,0.649,0.534
q8-hi 0.296,0.355,0.378,0.416,0.348,0.652,0.529
q8 0.299,0.354,0.378,0.414,0.346,0.650,0.535
q6-hi 0.301,0.356,0.378,0.415,0.350,0.651,0.541
q6 0.300,0.367,0.378,0.416,0.344,0.647,0.524
mxfp4 0.286,0.364,0.609,0.404,0.316,0.626,0.531
Quant Perplexity Peak memory
mxfp8 6.611 ยฑ 0.049 7.65 GB
mxfp4 7.455 ยฑ 0.057 6.33 GBDetailed metrics by model
nightmedia/Qwen3.5-0.8B-mxfp8-mlx
nightmedia/Qwen3.5-2B-mxfp8-mlx
nightmedia/Qwen3.5-4B-mxfp8-mlx
nightmedia/Qwen3.5-9B-mxfp8-mlx
https://huggingface.co/nightmedia/Qwen3.5-27B-Text
nightmedia/Qwen3.5-122B-A10B-Text-mxfp4-mlx
More metrics coming soon.
I am running these on my Mac, an M4Max with 128GB RAM. Some performance numbers like tokens/second reflect the performance on my box.
This post will be updated with every model that gets tested. The larger models take hours, the 27B a couple days, so it will be a long process.
-G
reacted to Csplk's post with ๐ 8 months ago
Post
2492
Was tinkering with a Daggr node generator script earlier today ( Csplk/DaggrGenerator )and started on a GUI for it for folks who are not comfy with writing code and like a GUI instead for something to motivate working on some Daggr stuff.
*Will have time later to keep working on it so donโt hesitate to comment with bugs or issues found if trying it out.*
Csplk/DaggrGenerator
Thanks @merve @ysharma @abidlabs and team daggr for making daggr :)
*Will have time later to keep working on it so donโt hesitate to comment with bugs or issues found if trying it out.*
Csplk/DaggrGenerator
Thanks @merve @ysharma @abidlabs and team daggr for making daggr :)
reacted to MonsterMMORPG's post with ๐๐ค 11 months ago
Post
2828
Qwen Image Models Training - 0 to Hero Level Tutorial - LoRA & Fine Tuning - Base & Edit Model - https://youtu.be/DPX3eBTuO_Y
This is a full comprehensive step-by-step tutorial for how to train Qwen Image models. This tutorial covers how to do LoRA training and full Fine-Tuning / DreamBooth training on Qwen Image models. It covers both the Qwen Image base model and the Qwen Image Edit Plus 2509 model. This tutorial is the product of 21 days of full R&D, costing over $800 in cloud services to find the best configurations for training. Furthermore, we have developed an amazing, ultra-easy-to-use Gradio app to use the legendary Kohya Musubi Tuner trainer with ease. You will be able to train locally on your Windows computer with GPUs with as little as 6 GB of VRAM for both LoRA and Fine-Tuning. Furthermore, I have shown how to train a character (person), a product (perfume) and a style (GTA5 artworks).
Tutorial Link : https://youtu.be/DPX3eBTuO_Y
This is a full comprehensive step-by-step tutorial for how to train Qwen Image models. This tutorial covers how to do LoRA training and full Fine-Tuning / DreamBooth training on Qwen Image models. It covers both the Qwen Image base model and the Qwen Image Edit Plus 2509 model. This tutorial is the product of 21 days of full R&D, costing over $800 in cloud services to find the best configurations for training. Furthermore, we have developed an amazing, ultra-easy-to-use Gradio app to use the legendary Kohya Musubi Tuner trainer with ease. You will be able to train locally on your Windows computer with GPUs with as little as 6 GB of VRAM for both LoRA and Fine-Tuning. Furthermore, I have shown how to train a character (person), a product (perfume) and a style (GTA5 artworks).
Tutorial Link : https://youtu.be/DPX3eBTuO_Y
reacted to flozi00's post with โค๏ธ 11 months ago
Post
3247
Some weeks ago, i've just decide its time to leave LinkedIn for me.
It got silent around my open source activities the last year, so i thought something has to change.
That's why my focus will move to share experiences and insights about hardware, drivers, kernels and linux. I won't post about how to use models, built agents or do prompting. I want to share about some deeper layers the actual hypes are built on.
I will start posting summarizations of my articles here on the hub.
English version:
https://flozi.net/en
German translated version:
https://flozi.net/de
Feel free to reach me if you want to read something specific.
It got silent around my open source activities the last year, so i thought something has to change.
That's why my focus will move to share experiences and insights about hardware, drivers, kernels and linux. I won't post about how to use models, built agents or do prompting. I want to share about some deeper layers the actual hypes are built on.
I will start posting summarizations of my articles here on the hub.
English version:
https://flozi.net/en
German translated version:
https://flozi.net/de
Feel free to reach me if you want to read something specific.
posted an update 11 months ago
Post
11603
Why I think local, open-source models will eventually win.
The most useful AI applications are moving toward multi-turn agentic behavior: systems that take hundreds or even thousands of iterative steps to complete a task, e.g. Claude Code, computer-control agents that click, type, and test repeatedly.
In these cases, the power of the model is not how smart it is per token, but in how quickly it can interact with its environment and tools across many steps. In that regime, model quality becomes secondary to latency.
An open-source model that can call tools quickly, check that the right thing was clicked, or verify that a code change actually passes tests can easily outperform a slightly โsmarterโ closed model that has to make remote API calls for every move.
Eventually, the balance tips: it becomes impractical for an agent to rely on remote inference for every micro-action. Just as no one would tolerate a keyboard that required a network request per keystroke, users wonโt accept agent workflows bottlenecked by latency. All devices will ship with local, open-source models that are โgood enoughโ and the expectation will shift toward everything running locally. Itโll happen sooner than most people think.
The most useful AI applications are moving toward multi-turn agentic behavior: systems that take hundreds or even thousands of iterative steps to complete a task, e.g. Claude Code, computer-control agents that click, type, and test repeatedly.
In these cases, the power of the model is not how smart it is per token, but in how quickly it can interact with its environment and tools across many steps. In that regime, model quality becomes secondary to latency.
An open-source model that can call tools quickly, check that the right thing was clicked, or verify that a code change actually passes tests can easily outperform a slightly โsmarterโ closed model that has to make remote API calls for every move.
Eventually, the balance tips: it becomes impractical for an agent to rely on remote inference for every micro-action. Just as no one would tolerate a keyboard that required a network request per keystroke, users wonโt accept agent workflows bottlenecked by latency. All devices will ship with local, open-source models that are โgood enoughโ and the expectation will shift toward everything running locally. Itโll happen sooner than most people think.
reacted to Severian's post with ๐ฅ๐ 12 months ago
Post
3281
MLX port of BDH (Baby Dragon Hatchling) is up!
Iโve ported the BDH ( https://github.com/pathwaycom/bdh ) model to MLX for Apple Silicon. Itโs a faithful conversion of the PyTorch version: same math, same architecture (byte-level vocab, shared weights across layers, ReLU sparsity, RoPE attention with Q=K), with MLX-friendly APIs and a detailed README explaining the few API-level differences and why results are equivalent.
Code, docs, and training script are ready to use. You may need to adjust the training script a bit to fit your own custom dataset. Only tested on M4 so far, but should work perfect for any M1/M2/M3 users out there.
Iโm currently training this MLX build on my Internal Knowledge Map (IKM) dataset Severian/Internal-Knowledge-Map
Trainingโs underway; expect a day or so before I publish weights. When itโs done, Iโll upload the checkpoint to Hugging Face for anyone to test.
Repo: https://github.com/severian42/BDH-MLX
HF model (coming soon): Severian/BDH-MLX
If you try it on your own data, feedback and PRs are welcome.
Iโve ported the BDH ( https://github.com/pathwaycom/bdh ) model to MLX for Apple Silicon. Itโs a faithful conversion of the PyTorch version: same math, same architecture (byte-level vocab, shared weights across layers, ReLU sparsity, RoPE attention with Q=K), with MLX-friendly APIs and a detailed README explaining the few API-level differences and why results are equivalent.
Code, docs, and training script are ready to use. You may need to adjust the training script a bit to fit your own custom dataset. Only tested on M4 so far, but should work perfect for any M1/M2/M3 users out there.
Iโm currently training this MLX build on my Internal Knowledge Map (IKM) dataset Severian/Internal-Knowledge-Map
Trainingโs underway; expect a day or so before I publish weights. When itโs done, Iโll upload the checkpoint to Hugging Face for anyone to test.
Repo: https://github.com/severian42/BDH-MLX
HF model (coming soon): Severian/BDH-MLX
If you try it on your own data, feedback and PRs are welcome.
Nice!!!
reacted to piercus's post with ๐ 12 months ago
Post
2982
We've just forked LBM to reproduce the LBM eraser results
Our fork : https://github.com/finegrain-ai/LBM
LBM paper: LBM: Latent Bridge Matching for Fast Image-to-Image Translation (2503.07535)
LBM relighting demo : jasperai/LBM_relighting
Our fork : https://github.com/finegrain-ai/LBM
LBM paper: LBM: Latent Bridge Matching for Fast Image-to-Image Translation (2503.07535)
LBM relighting demo : jasperai/LBM_relighting
posted an update about 1 year ago
Post
1633
What other features would you like to see on the Trackio Dashboard? ( gradio-templates/trackio-dashboard)
reacted to merve's post with ๐ฅ about 1 year ago
Post
7346
large AI labs open-sourced a ton of models last week ๐ฅ
here's few picks, find even more here merve/sep-16-releases-68d13ea4c547f02f95842f05 ๐ค
> IBM released a new Docling model with 258M params based on Granite (A2.0) ๐ ibm-granite/granite-docling-258M
> Xiaomi released 7B audio LM with base and instruct variants (MIT) XiaomiMiMo/mimo-audio-68cc7202692c27dae881cce0
> DecartAI released Lucy Edit, open Nano Banana ๐ (NC) decart-ai/Lucy-Edit-Dev
> OpenGVLab released a family of agentic computer use models (3B/7B/32B) with the dataset ๐ป OpenGVLab/scalecua-68c912cf56f7ff4c8e034003
> Meituan Longcat released thinking version of LongCat-Flash ๐ญ meituan-longcat/LongCat-Flash-Thinking
here's few picks, find even more here merve/sep-16-releases-68d13ea4c547f02f95842f05 ๐ค
> IBM released a new Docling model with 258M params based on Granite (A2.0) ๐ ibm-granite/granite-docling-258M
> Xiaomi released 7B audio LM with base and instruct variants (MIT) XiaomiMiMo/mimo-audio-68cc7202692c27dae881cce0
> DecartAI released Lucy Edit, open Nano Banana ๐ (NC) decart-ai/Lucy-Edit-Dev
> OpenGVLab released a family of agentic computer use models (3B/7B/32B) with the dataset ๐ป OpenGVLab/scalecua-68c912cf56f7ff4c8e034003
> Meituan Longcat released thinking version of LongCat-Flash ๐ญ meituan-longcat/LongCat-Flash-Thinking
reacted to meg's post with โค๏ธ about 1 year ago
Post
2997
๐ค As AI-generated content is shared in movies/TV/across the web, there's one simple low-hanging fruit ๐ to help know what's real: Visible watermarks. With the Gradio team, I've made sure it's trivially easy to add this disclosure to images, video, chatbot text. See how: https://huggingface.co/blog/watermarking-with-gradio
Thanks to the code collab in particular from @abidlabs and Yuvraj Sharma.
Thanks to the code collab in particular from @abidlabs and Yuvraj Sharma.
reacted to hba123's post with ๐ about 1 year ago
Post
2891
We have written a fun little blog on how you can do robotics with Ark and in Python. We also give you some examples of how OpenAI Gym can become hardware-grounded and how easy it is to do so:
Check it out: https://huggingface.co/blog/hba123/ark
Check it out: https://huggingface.co/blog/hba123/ark
posted an update over 1 year ago
Post
4519
The Gradio x Agents x MCP hackathon keeps growing! We now have more $1,000,000 in credit for participants and and >$16,000 in cash prizes for winners.
We've kept registration open until the end of this week, so join and let's build cool stuff together as a community: https://huggingface.co/spaces/ysharma/gradio-hackathon-registration-2025
We've kept registration open until the end of this week, so join and let's build cool stuff together as a community: https://huggingface.co/spaces/ysharma/gradio-hackathon-registration-2025