It’s about GPUs. Got me thinking maybe this would also work for FP128, or CPUs (or GPUs).
It’s not simply double-double scheme (is that’s Dekker’s?). It doesn’t supports scalars (or well), pays off with large enough matrix. Probably Julia can’t do Float64 faster on CPUs, using Float32, I suppose all CPUs have as many units for, or half as fast bandwidth wise.