Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🔄
In a Training Loop
36.4
TFLOPS
Giles Thomas
gpjt
3
Follow
Arp25's profile picture
AnthonyPa57's profile picture
fauverism's profile picture
16 followers
·
27 following
https://www.gilesthomas.com/
gpjt
gpjt
gilesthomas
gilesthomas.com
AI & ML interests
Doing my best to speedrun 20 years of AI research. YMMV
Recent Activity
updated
a model
about 3 hours ago
gpjt/jax-with-mha-bias-openwebtext
published
a model
about 3 hours ago
gpjt/jax-with-mha-bias-openwebtext
new
activity
1 day ago
gpjt/jax-with-mha-bias-fw-fwedu-5050:
Factorized embeddings and non-embedding parameter allocation for GPT-2
View all activity
Organizations
None yet
gpjt
's models
42
Sort:Â Recently updated
gpjt/jax-with-mha-bias-openwebtext
Text Generation
•
0.2B
•
Updated
about 3 hours ago
gpjt/jax-with-mha-bias-fw-fwedu-simplewiki
Text Generation
•
0.2B
•
Updated
2 days ago
•
535
gpjt/jax-with-mha-bias-fw-fwedu-5050
Text Generation
•
0.2B
•
Updated
2 days ago
•
155
•
1
gpjt/1xrtx3090-moe-1
Text Generation
•
0.5B
•
Updated
3 days ago
•
493
gpjt/jax-with-mha-bias-fineweb-edu
Text Generation
•
0.2B
•
Updated
3 days ago
•
264
gpjt/jax-with-mha-bias-larger-chinchilla-1
Text Generation
•
0.3B
•
Updated
14 days ago
•
284
gpjt/jax-with-mha-bias-larger-chinchilla-2
Text Generation
•
0.2B
•
Updated
14 days ago
•
274
gpjt/jax-with-mha-bias-no-dropout-2-epoch
Text Generation
•
0.2B
•
Updated
14 days ago
•
234
gpjt/jax-with-mha-bias-no-dropout-extended
Text Generation
•
0.2B
•
Updated
14 days ago
•
256
gpjt/jax-no-mha-bias-no-dropout
Text Generation
•
0.2B
•
Updated
14 days ago
•
329
gpjt/jax-with-mha-bias-no-dropout
Text Generation
•
0.2B
•
Updated
14 days ago
•
285
gpjt/jax-no-mha-bias-with-dropout
Text Generation
•
0.2B
•
Updated
14 days ago
•
257
gpjt/1xrtx3090-stacked-interventions
Text Generation
•
0.2B
•
Updated
Apr 15
•
16
gpjt/1xrtx3090-baseline
Text Generation
•
0.2B
•
Updated
Apr 14
•
16
gpjt/8xa100m40-stacked-interventions-3
Text Generation
•
0.2B
•
Updated
Apr 9
•
18
gpjt/8xa100m40-stacked-interventions-2
Text Generation
•
0.2B
•
Updated
Apr 8
•
18
gpjt/8xa100m40-stacked-interventions-1
Text Generation
•
0.2B
•
Updated
Apr 8
•
15
gpjt/8xa100m40-baseline-8
Text Generation
•
0.2B
•
Updated
Apr 8
•
20
gpjt/8xa100m40-baseline-7
Text Generation
•
0.2B
•
Updated
Apr 8
•
16
gpjt/8xa100m40-baseline-6
Text Generation
•
0.2B
•
Updated
Apr 8
•
19
gpjt/8xa100m40-baseline-4
Text Generation
•
0.2B
•
Updated
Apr 8
•
16
gpjt/8xa100m40-baseline-5
Text Generation
•
0.2B
•
Updated
Apr 8
•
15
gpjt/8xa100m40-baseline-3
Text Generation
•
0.2B
•
Updated
Apr 8
•
19
gpjt/8xa100m40-baseline-2
Text Generation
•
0.2B
•
Updated
Apr 8
•
17
gpjt/8xa100m80-no-amp
Text Generation
•
0.2B
•
Updated
Apr 3
•
19
gpjt/8xa100m40-weight-decay-cerebras
Text Generation
•
0.2B
•
Updated
Mar 24
•
18
gpjt/8xa100m40-weight-decay-gpt2
Text Generation
•
0.2B
•
Updated
Mar 24
•
16
gpjt/8xa100m40-qkv-bias
Text Generation
•
0.2B
•
Updated
Mar 24
•
19
gpjt/8xa100m40-schedule-learning-rate
Text Generation
•
0.2B
•
Updated
Mar 24
•
18
gpjt/8xa100m40-remove-dropout
Text Generation
•
0.2B
•
Updated
Mar 24
•
17
Previous
1
2
Next