Sharing on Mastodon:
Cutting RAG inference costs 6x starts with deciding what never reaches the LLM
Save
Home
About