Fast AI is more than tokens per second.
The first visible response, total completion time, and the quality of the finished task measure different things.

Speed is a product experience as well as a model metric. A response that starts quickly can still take a long time to finish. A shorter completion can be unhelpful if it does not solve the task.
Measure the moments that matter
Record how long the user waits for useful feedback and how long the complete operation takes. Test under realistic network conditions with representative inputs. Compare the same task and output requirements when assessing providers.
Use streaming deliberately
Streaming delivers a response in incremental events. Anthropic's documentation shows that streams can include content, completion signals, and errors. Treat progress as progress: receiving some output does not mean the operation succeeded.
Design the waiting state
Show a clear status, preserve the user's input after an error, and explain what a retry will do. For generated code, validate the complete result before treating it as a usable project. A responsive interface should also remain understandable when the connection fails.