>1B tokens/minute/GPU by combining query planner and inference engine

(modal.com)

7 points | by charles_irl 13 hours ago

2 comments