I've been exploring how we can use vision-language models (VLMs) to optimize object classification tasks, drawing inspiration from workflows like Rohit Singh's demonstrations of combining image and text context for robust spatial reasoning.
Before I start testing, I wanted to ask one question:
Usage and quota impact: Feeding high-resolution images into OpenAI model uses up tokens quickly. I'm curious how hard this will hit our limits. Has anyone tested how quickly a workflow like this consumes a standard five-hour testing allotment?