Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🏗️
Building on HF
9.2
TFLOPS
Sergio Paniego
PRO
sergiopaniego
245
208
136
Follow
pgpt19's profile picture
h9052300's profile picture
manikantthakur's profile picture
2,008 followers
·
135 following
https://sergiopaniego.github.io/
sergiopaniego
sergiopaniego
sergio-paniego-blanco
AI & ML interests
None yet
Recent Activity
liked
a dataset
about 3 hours ago
FineEnvs/HF_ML_Tasksmith
posted
an
update
about 3 hours ago
Repo2RLEnv just shipped TaskSmith + 50 high quality RL envs generated from HF repos 🔨 TaskSmith is a specialized harness that turns a merged PR into a verified RL env the envs come from HF repos (Transformers, TRL, PEFT, Accelerate, Diffusers), shipped as Harbor tasks you can eval or train on > code: github.com/huggingface/Repo2RLEnv > dataset: huggingface.co/datasets/FineEnvs/HF_ML_Tasksmith
new
activity
about 7 hours ago
agents-course/Unit4-Final-Certificate:
Certificate page: Sign in button not showing up. Shows Remi Testa
View all activity
Organizations
sergiopaniego
's models
141
Sort: Recently updated
sergiopaniego/qwen3-1.7b-mbpp-grpo
Text Generation
•
2B
•
Updated
26 days ago
•
477
sergiopaniego/rl-envs-youtube-livestream-4-scripts
Updated
26 days ago
sergiopaniego/qwen3-1.7b-wordle-grpo
Text Generation
•
2B
•
Updated
Aug 19
•
59
sergiopaniego/Qwen3-4B-claude-code-local-grpo
Text Generation
•
4B
•
Updated
Jul 31
•
12
sergiopaniego/Qwen3-4B-claude-code-deepcoder-grpo
Text Generation
•
4B
•
Updated
Jul 31
•
13
sergiopaniego/pelican-svg-grpo-Qwen3-1.7B-judged
Text Generation
•
2B
•
Updated
Jul 29
•
10
sergiopaniego/pelican-svg-grpo-Qwen3-1.7B
Text Generation
•
2B
•
Updated
Jul 29
•
12
sergiopaniego/Qwen3-8B-opencode-deepcoder-grpo
Text Generation
•
8B
•
Updated
Jul 28
•
47
•
1
sergiopaniego/grpo-youtube-livestream-3-scripts
Reinforcement Learning
•
Updated
Jul 27
•
2
sergiopaniego/Qwen3.5-4B-sdpo-math-hints
Updated
Jul 10
•
1
sergiopaniego/Qwen3.5-4B-sdpo-math-gold
Updated
Jul 10
sergiopaniego/Qwen3.5-4B-sdpo-math-baseline
Updated
Jul 10
sergiopaniego/sdpo-hints
Updated
Jul 10
sergiopaniego/pi-mono-youtube-livestream-2-scripts
Updated
Jul 6
•
2
sergiopaniego/gemma-4-E2B-offpolicy-kd-lr1e4
Updated
Jul 6
sergiopaniego/gemma-4-E2B-offpolicy-kd-lr2e4
Updated
Jul 6
sergiopaniego/gemma-4-E2B-offpolicy-kd-lr5e5
Updated
Jul 6
sergiopaniego/qwen3-0.6b-pimono-gkd-lr2e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
12
sergiopaniego/qwen3-0.6b-pimono-gkd-lr1e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
16
sergiopaniego/qwen3-0.6b-pimono-gkd-lr5e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
16
sergiopaniego/qwen3-0.6b-pimono-logit-kd-lr1e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
11
sergiopaniego/qwen3-0.6b-pimono-logit-kd-lr2e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
12
sergiopaniego/qwen3-0.6b-pimono-logit-kd-lr5e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
7
sergiopaniego/Qwen2.5-0.5B-Instruct-text-to-sql-qlora
Updated
Jun 15
sergiopaniego/browsergym-grpo-functiongemma-270m-it
Text Generation
•
0.3B
•
Updated
May 29
•
30
•
2
sergiopaniego/qwen3-grpo-requests
Updated
May 19
sergiopaniego/reasoning-gym-chain-sum-Qwen3-1.7B-sft
Text Generation
•
2B
•
Updated
May 4
•
12
sergiopaniego/reasoning-gym-chain-sum-Qwen3-1.7B
Text Generation
•
2B
•
Updated
Apr 28
•
17
sergiopaniego/carla-vlm-gemma-test
Updated
Apr 15
sergiopaniego/carla-vlm-qwen35-test
Updated
Apr 13
Previous
1
2
3
...
5
Next