I recall you sharing a research paper a few months ago regarding distributed training or inference.
@Palli FYI, I just found the papers poasted by you on this topic. They are about DisTrO, SuperScaler, DeepSpeed, Megatron, Alpa, DoReMi. Mostly in your posts (number 11/8/4) of the original thread.