Multimodal APIs are usually documented one modality at a time — here is image input, here is audio, here is streaming text. What the docs rarely cover is how to hold them in one coherent session, which is what an actual application needs.




Parts, not endpoints


The mental shift that made this tractable: stop thinking in endpoints, start...