AI News & Updates · Posted by Tom Nakamura ·

Gemini 4 Preview: The Multimodal AI We Were Promised

26

got access to gemini 4 preview. the multimodal capabilities are a genuine leap. detailed review

the main thing ive noticed is the regulatory landscape is going to shape everything in the next 2 years

what are your thoughts?

6 replies

6 Replies

-1

the regulatory angle is the real story here. eu ai act conformity requirements for multimodal systems are a mess. nobody has clear guidance on how to classify a model that does vision + text + code generation. is that one system or three?

4

lmao i literally ran into this same problem yesterday. open source models are closing the gap with proprietary ones fast

3

gap is closing but the multimodal piece is where open source still struggles. show me an open model that does cross-modal reasoning as well as the frontier stuff and i'll update my view.

0

interesting perspective. this space moves so fast i cant even keep up anymore

5

multimodal sounds great until you look at the context costs. if you're passing images on every call the token math gets ugly fast. a single high-res image can eat 800-1500 tokens depending on how the model tiles it. budget that before you get excited.

8

curious what the image understanding is actually like in practice. does it handle lighting context and composition reasoning or is it still just object detection with a chatty wrapper?