2 points | by soltanov 4 hours ago ago
1 comments
Been waiting for MTP in llamacpp for a about a week now so I can speed up 3.8:flash-next but this might be just what I'm looking for. :)
Update: Looks good so far
Flash-Next on llama.cpp with no MTP: 28.8t/s decode, ~620t/s prefill
on llama.cpp with MTP (froze twice): 44.1t/s decode, ~593t/s prefill
on TensorFold + MTP: 61.0t/s decode, 2,455t/s prefill
That's about 2.1 to 2.4x on decode and about 4x on prefill (!!)
Need to do some long context testing but Im stoked.
Been waiting for MTP in llamacpp for a about a week now so I can speed up 3.8:flash-next but this might be just what I'm looking for. :)
Update: Looks good so far
Flash-Next on llama.cpp with no MTP: 28.8t/s decode, ~620t/s prefill
on llama.cpp with MTP (froze twice): 44.1t/s decode, ~593t/s prefill
on TensorFold + MTP: 61.0t/s decode, 2,455t/s prefill
That's about 2.1 to 2.4x on decode and about 4x on prefill (!!)
Need to do some long context testing but Im stoked.