while true, prefill performance is still a huge roadblock
while true, prefill performance is still a huge roadblock 50k tokens of input still unusable on local, that was not the case on gpt4 cloud api even if output tok/s matched the UX lag tax is kind of a big deterrent for the time being