Braid
Training data with a clear source, clear usage rights, and repeatable test sets.
Machine-learning research · Richmond, BC
Cenetex Lab trains small language models. Their size lets us see what helps: better data, better training order, or better tests.
settings learned by the same small research model
Why we start small
Large models can hide what caused an improvement. Zero is small on purpose. We keep the model the same, change one thing, and test whether it really helped.
If an idea works, we can try it on a larger model. If it fails, we learn why before spending much more money.
The research system
Training data with a clear source, clear usage rights, and repeatable test sets.
The same small model used to test better data, better training order, and better ways to learn.
Tests that check facts and relationships, not only whether the writing sounds good.
A clear improvement
In C3.1, we mixed the training tasks instead of teaching them in separate blocks. Test loss fell from 1.7623 to 1.5666. Everything else stayed the same.
Zero lineage
We built a way to split text that loses nothing, used facts we can trace, and made training repeatable.
After reading 11.9 million pieces of text, the model began to show a basic sense of language and meaning.
The model improved, but often changed its answer when we reversed the order of two choices.
We now show both orders in the same learning step. The final result is not ready yet.
Current experiment · C3.3
The model now sees a pair and its reversed version at the same time. We saved its progress after 7,000 of 9,442 training steps. We will not call it a success until the final tests are complete.
What must improve
Build with Cenetex
We work with teams that want measurable machine learning—not a demo that only looks intelligent.
Talk to Cenetex →