TTFT / TPS
Why do some AI tools take forever to show the first word while others answer right away?
Why do some AI tools take forever to show the first word while others answer right away?
These times and speeds are illustrative; they only compare how the metrics relate.
B starts speaking sooner, but generates more slowly afterward, so its full answer finishes later.
Wait for TTFT: TTFT runs from send to the first text token.
Then watch TPS: TPS is tokens per second after generation begins, not transactions per second.
Then measure total time: A longer answer can still take longer to finish even when it begins quickly.
Measure TTFT, TPS after output begins, and total completion separately under the same model settings and answer target. Do not replace full comparison with feels instant, and record network, queue, and generation phases.
Build in 3D. Put your computer to work.