For three years the default assumption was that serious inference happened somewhere else, on hardware you rented by the token. That assumption is now wrong for a widening band of tasks.
Where the line sits
Classification, extraction, rewriting, and most retrieval-augmented question answering run acceptably on a machine you already own. Long-horizon reasoning still does not.
The interesting part is not the benchmark. It is that the tasks below the line no longer generate a log entry on someone else’s server, which changes what you are willing to feed them.