Integrated graphics handled these local AI models better than expected.
I stopped using my RTX GPU for every LLM.