Building a 3D Paris Gallery with AI Agents and Callable Spaces
A single coding agent just built a full 3D gallery of Paris monuments without touching a single 3D tool, proving that UI endpoints are the new npm packages....

I watched a coding agent spin up a fully rendered 3D gallery of Paris monuments the other day, and I didn't see a single line of traditional 3D modeling code. No Blender files. No messy manual asset pipelines. The agent simply chained two distinct Hugging Face Spaces together, turning a text prompt into crisp image specimens, and then transforming those images into native 3D Gaussian splats entirely on its own. It's wild.
Mitchell Hashimoto recently coined the 'building — to be fair — block economy' thesis to describe how software is made now. This Instead of wrestling with giant, brittle monoliths, we're assembling small, well-documented components! From what I can tell, why — or more precisely, does this matter? Strictly, while most people apply this idea to code libraries, watching an LLM orchestrate live multimedia endpoints changes everything. The heavy lifting of AI has never really been the model weights. It has always been the miserable plumbing of SDKs. Incompatible input formats, and GPU orchestration. It has always been the miserable plumbing of SDKs, incompatible input formats, and GPU orchestration.
Enter the understated brilliance of machine learning infrastructure catching up to reality. Every Gradio deployment on the hub now quietly serves an agents.md file, essentially handing a fully formed API manual directly to any autonomous loop that asks for it. An agent curls a single URL, reads the exact schema, handles the auth handshake, and immediately starts piping outputs from a text-to-image generator straight into a 3D reconstruction pipeline. No custom client libraries required.

This is what real use looks like. We're moving past the era where building a multimedia use meant spending three weeks wrestling with client-side adding and docs hell. When every model on the internet exposes a standardized, machine-readable interface, human developers stop gluing APIs together manually. We step up a layer, define the intent, and let the agent navigate the ecosystem for us.
If you are still writing custom wrapper code for every single model inference call you make, you are doing work that machines are now fully equipped to handle. The stack has flattened. The components are live. Go build something weird.






