Datasets and an interactive explorer connecting Japan, Iceland, mathematics, technology, philosophy, Asian culture and Christian ethics.
Guilherme Monteiro
guicybercode
·
AI & ML interests
smol-llms, llm-training, chinese-history, brazilian-history, production-ml, terminal-ui, data-curation, edge-ai
Recent Activity
updated a collection 13 days ago
Culture, Mathematics, Technology & Ethics updated a Space 13 days ago
guicybercode/culture-compass published a Space 14 days ago
guicybercode/culture-compassOrganizations
None yet
LLM Training from Scratch
Resources, datasets and base models for training small LLMs on consumer hardware. Focus on pre-training pipelines and data curation.
Smol LLMs for History
Small language models (≤3B params) that can be fine-tuned or run locally for history-focused tasks. Curated for training on Chinese and Brazilian hist
-
HuggingFaceTB/SmolLM2-135M
Text Generation • 0.1B • Updated • 2.4M • 231 -
HuggingFaceTB/SmolLM2-135M-Instruct
Text Generation • 0.1B • Updated • 1.35M • 414 -
HuggingFaceTB/SmolLM2-360M
Text Generation • 0.4B • Updated • 450k • 128 -
HuggingFaceTB/SmolLM2-360M-Instruct
Text Generation • 0.4B • Updated • 304k • 220
Papers I'm Reading
Key papers on small LLMs, data-centric training, and multilingual NLP.
-
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Paper • 2406.17557 • Published • 106 -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Paper • 2502.02737 • Published • 260 -
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Paper • 2506.05209 • Published • 65
Multilingual models
Multilingual models featuring Portuguese, Chinese, and other languages essential for cross-cultural historical research. Curated for fine-tuning on Ch
-
Qwen/Qwen2.5-0.5B
Text Generation • 0.5B • Updated • 1.66M • • 444 -
Qwen/Qwen2.5-0.5B-Instruct
Text Generation • 0.5B • Updated • 7.23M • • 623 -
Qwen/Qwen2.5-1.5B-Instruct
Text Generation • 2B • Updated • 7.27M • • 824 -
Qwen/Qwen2.5-1.5B-Instruct-GGUF
Text Generation • 2B • Updated • 195k • 150
Culture, Mathematics, Technology & Ethics
Datasets and an interactive explorer connecting Japan, Iceland, mathematics, technology, philosophy, Asian culture and Christian ethics.
Papers I'm Reading
Key papers on small LLMs, data-centric training, and multilingual NLP.
-
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Paper • 2406.17557 • Published • 106 -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Paper • 2502.02737 • Published • 260 -
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Paper • 2506.05209 • Published • 65
LLM Training from Scratch
Resources, datasets and base models for training small LLMs on consumer hardware. Focus on pre-training pipelines and data curation.
Multilingual models
Multilingual models featuring Portuguese, Chinese, and other languages essential for cross-cultural historical research. Curated for fine-tuning on Ch
-
Qwen/Qwen2.5-0.5B
Text Generation • 0.5B • Updated • 1.66M • • 444 -
Qwen/Qwen2.5-0.5B-Instruct
Text Generation • 0.5B • Updated • 7.23M • • 623 -
Qwen/Qwen2.5-1.5B-Instruct
Text Generation • 2B • Updated • 7.27M • • 824 -
Qwen/Qwen2.5-1.5B-Instruct-GGUF
Text Generation • 2B • Updated • 195k • 150
Smol LLMs for History
Small language models (≤3B params) that can be fine-tuned or run locally for history-focused tasks. Curated for training on Chinese and Brazilian hist
-
HuggingFaceTB/SmolLM2-135M
Text Generation • 0.1B • Updated • 2.4M • 231 -
HuggingFaceTB/SmolLM2-135M-Instruct
Text Generation • 0.1B • Updated • 1.35M • 414 -
HuggingFaceTB/SmolLM2-360M
Text Generation • 0.4B • Updated • 450k • 128 -
HuggingFaceTB/SmolLM2-360M-Instruct
Text Generation • 0.4B • Updated • 304k • 220