https://howtonotcode.com/story/403-glm-47-hits-real-time-speeds-on-cerebras-for-coding-and-agent-workflows