How ternary weights and BitNet b1.58 matched full-precision language models at a fraction of the cost