Post
251
I've been gradually recaptioning aged text-to-image datasets with better vision models! These datasets are also repackaged into the more modern webshart format (https://github.com/bghira/webshart) which includes built-in aspect bucketing and caption delivery.
The first two datasets are ready for use!
- webshart/terminusresearch-photo-anatomy
- webshart/terminusresearch-photo-aesthetics
"anatomy" is a bunch of human-centric images containing people holding or otherwise interacting with objects or positioned in complex ways.
"aesthetics" is a collection of visually striking images - high contrast, diverse colouration, and cinematic framing (among other factors).
These two datasets from 2023 were recaptioned with GLM 5.3 Flash via Ollama Cloud and Zhipu AI APIs, with a smaller portion run over 4x H100 with GLM 5.3 Flash in W4A16 precision - these outputs were checked by hand for quality, and the 4bit run was continued at a batch size of 64.
What datasets would you like to see recaptioned next? A better CC12M is on its way!
The first two datasets are ready for use!
- webshart/terminusresearch-photo-anatomy
- webshart/terminusresearch-photo-aesthetics
"anatomy" is a bunch of human-centric images containing people holding or otherwise interacting with objects or positioned in complex ways.
"aesthetics" is a collection of visually striking images - high contrast, diverse colouration, and cinematic framing (among other factors).
These two datasets from 2023 were recaptioned with GLM 5.3 Flash via Ollama Cloud and Zhipu AI APIs, with a smaller portion run over 4x H100 with GLM 5.3 Flash in W4A16 precision - these outputs were checked by hand for quality, and the 4bit run was continued at a batch size of 64.
What datasets would you like to see recaptioned next? A better CC12M is on its way!