BIMgent: Generating Building Models via Computer-use Agents
Building Information Modeling (BIM) authoring in the Architecture, Engineering, and Construction (AEC) sector entails navigating complex professional software, demanding significant manual effort and domain expertise. While current automation approaches rely heavily on rigid Application Programming Interfaces (APIs), they often detach agents from human-centric visual workflows, with API-specific implementations further limiting their generalizability across software platforms. In this work, we introduce BIMgent, a multimodal LLM-based agentic framework designed to perform autonomous BIM authoring directly via Graphical User Interface (GUI) interactions. To mitigate the challenges of high visual noise and long-horizon dependencies inherent in BIM authoring software, BIMgent incorporates a hierarchical planning mechanism guided by software documentation and employs dynamic GUI grounding to precisely navigate complex user interfaces. We evaluate our framework on a benchmark derived from real-world residential building designs. Experimental results demonstrate that BIMgent achieves a 77.8\% end-to-end success rate in correctly generating semantically rich BIM models from input floorplans, significantly outperforming the leading generalist computer-use models, including Claude-Sonnet-4.5 (2.2\%) and GPT-5.4 (11.1\%). By bridging the gap between abstract design inputs and concrete modeling actions without API constraints, BIMgent presents a viable and generalizable path toward autonomous architectural modeling assistants.