3PO family methods trained on DapoMath-17k using Olmo3-IVON-SFT-7B and Qwen2.5Math-IVON-SFT-7B
BayesRL
non-profit
AI & ML interests
None defined yet.
Recent Activity
View all activity
A collection of three models trained on the Nemotron Post Training Dataset for reasoning tasks with IVON
-
BayesRL/Llama3.1-IVON-SFT-8B
Text Generation ⢠8B ⢠Updated ⢠6.34k -
BayesRL/Qwen2.5Math-IVON-SFT-7B
Text Generation ⢠8B ⢠Updated ⢠926 -
BayesRL/Olmo3-IVON-SFT-7B
Text Generation ⢠7B ⢠Updated ⢠626 -
Parameter Exploration for RLVR via Variational Learning
Paper ⢠2608.09805 ⢠Published ⢠8
3PO family methods trained on DapoMath-17k using Olmo3-IVON-SFT-7B and Qwen2.5Math-IVON-SFT-7B
A collection of three models trained on the Nemotron Post Training Dataset for reasoning tasks with IVON
-
BayesRL/Llama3.1-IVON-SFT-8B
Text Generation ⢠8B ⢠Updated ⢠6.34k -
BayesRL/Qwen2.5Math-IVON-SFT-7B
Text Generation ⢠8B ⢠Updated ⢠926 -
BayesRL/Olmo3-IVON-SFT-7B
Text Generation ⢠7B ⢠Updated ⢠626 -
Parameter Exploration for RLVR via Variational Learning
Paper ⢠2608.09805 ⢠Published ⢠8