Local LLMs as System Tooling: Running many Sub-3B Models on Edge and Budget Hardware
You need an AI assistant that runs on your own machine. The cloud options are good, but they are not yours. You want something that actually does things — reads your screen, controls your smart home, answers questions in Finnish or English, and never sends your data to a server you don’t control. The local LLM trend is about as practical as it sounds. It works on mid-range hardware if you pick the right model size. A sub-3B parameter model fits comfortably on consumer gear and handles everyday tasks without needing a dedicated GPU or a rack of servers. The trade-off versus cloud options is that the response times are slower, but your parents’ conversations do not become training data for a future model owned by a third party. ...