Which is the best open model for [Julia] coding? And practical, and still reasonably good

I recall you sharing a research paper a few months ago regarding distributed training or inference.

@Palli FYI, I just found the papers poasted by you on this topic. They are about DisTrO, SuperScaler, DeepSpeed, Megatron, Alpa, DoReMi. Mostly in your posts (number 11/8/4) of the original thread.