Gemini 4 Preview: The Multimodal AI We Were Promised
got access to gemini 4 preview. the multimodal capabilities are a genuine leap. detailed review
the main thing ive noticed is the regulatory landscape is going to shape everything in the next 2 years
what are your thoughts?
6 Replies
Join the discussion.
Log In to Replylmao i literally ran into this same problem yesterday. open source models are closing the gap with proprietary ones fast
gap is closing but the multimodal piece is where open source still struggles. show me an open model that does cross-modal reasoning as well as the frontier stuff and i'll update my view.
interesting perspective. this space moves so fast i cant even keep up anymore
multimodal sounds great until you look at the context costs. if you're passing images on every call the token math gets ugly fast. a single high-res image can eat 800-1500 tokens depending on how the model tiles it. budget that before you get excited.
curious what the image understanding is actually like in practice. does it handle lighting context and composition reasoning or is it still just object detection with a chatty wrapper?
the regulatory angle is the real story here. eu ai act conformity requirements for multimodal systems are a mess. nobody has clear guidance on how to classify a model that does vision + text + code generation. is that one system or three?