Skip to content

The vision-based shift: Transforming construction safety with AI

The vision-based shift: Transforming construction safety with AI

There is a camera pointed at almost every corner of a modern Singapore construction site. Drones are flying scheduled routes overhead. IoT sensors on entry points track who comes and goes. Site conditions are continuously captured and uploaded to project management platforms. All that visual data. All those AI cameras. But the accidents kept happening in places where the cameras were pointed directly at. The question worth sitting with in 2026 is not whether Singapore’s construction sites are monitored enough. They are. The question is whether any of that monitoring actually understands what it is seeing. The construction industry’s safety problem is not a lack of cameras. It is a lack of comprehension. And that is a problem that only AI Agents with genuine vision can solve. What every construction site camera has always been missing A visit to any large construction project in Singapore will show that the technological monitoring infrastructure is impressive. From fixed AI cameras with video analytics, LiDAR-based systems, drone feeds capturing progress from above, access control sensors logging every worker movement, and IoT devices monitoring structural conditions, the data volume is enormous. The intelligence extracted from it, historically, however, has been almost none. Traditional video analytics, the software layer that has sat on top of this infrastructure for the past decade, was built to detect, not to understand. It was trained to recognise specific objects like a hard hat, a hi-vis vest, and a person crossing a danger zone. When those objects matched a predefined rule, it fired an alert. When they did not match, because lighting had changed, or the angle was different, or the hazard was one the system had never seen before, it either stayed silent or generated noise. Safety teams on large Singapore sites know this cycle intimately. The system fires hundreds of alerts per shift. Most are false positives. The team learns to filter. And in the filtering, somewhere in the volume, a real hazard goes unaddressed. Also Read: When AI agents start acting on our behalf, security gets more complicated The alert fatigue is so well-documented that it has its own name in the industry, and it has persisted for years precisely because no one has found a way to make these systems understand context rather than just detect objects. That is what changed when we gave construction sites in Singapore “eyes”. Vision language models: The technology that gave AI agents sight A Vision Language Model (VLM) is not an incremental improvement on computer vision. It is a different category of capability entirely. Where a traditional detection model draws a box around an object it recognises, a VLM reads a visual scene the way a trained human observer would, interpreting what is happening, what the relationship between elements in the frame means, and what risk that combination represents. When visual reasoning is embedded into AI agents deployed on construction sites, three things change in immediate and practical ways. The first is near-miss capture. For instance, the National Safety Council reports near-misses among the most underreported events in industrial safety. Workers do not always recognise them as near-misses, supervisors are not always present, and reporting friction means most events go unlogged. A visually intelligent agent does not depend on anyone reporting anything. It observes continuously, classifies events in real time, and builds an auditable record that reveals the precursor patterns to serious injury long before those injuries occur. The second is permit-to-work verification. Singapore’s high-risk worksites run on PTW systems designed to control access to areas where confined space entry, hot work, or work at height is underway. The persistent limitation of paper-based PTW is that it verifies the existence of an authorisation, not the state of the site. A VLM agent can do that visually: confirming that barriers are in place, energy sources isolated, rescue equipment present, before work begins. The permit and the physical site either agree or they do not. The third is temporal reasoning, which is comparing site imagery across time to determine whether a flagged hazard has been genuinely remediated. When a safety officer flags an unguarded edge, the agent later confirms from site imagery alone that the barrier has been properly installed. Each confirmation is timestamped, location-tagged, and added to the compliance record automatically. Together, these capabilities shift safety management to autonomous, continuous prevention. That is the shift MOM has been asking the construction sector to make. For decades, the construction industry built sophisticated systems to watch its sites and then did almost nothing with what those systems saw. We were generating visual intelligence at scale and discarding it at scale. VLM-powered agents are the first technology that actually closes that loop, and when you close it, the entire logic of how you manage safety on a site changes permanently. What this means for Singapore’s WSH ambitions MOM’s 2025 WSH Report marked a genuine milestone: Singapore’s workplace fatal injury rate fell to a record low of 0.96 per 100,000 workers, placing the country among the world’s safest working environments. The construction sector’s fatal and major injury rate dropped from 31.0 in 2024 to 26.3 in 2025, driven by stepped-up enforcement, two sector-wide safety time-outs, and stricter public-sector tender requirements. Also Read: Why AI agents will reshape customer journeys in Southeast Asia The progress is real. It also has a ceiling, and the industry is approaching it. The tools that got Singapore’s construction safety record to where it is today are running close to their limits of effectiveness. MOM’s own 2024 report acknowledged this, pointing explicitly to video surveillance systems and technology-driven hazard detection as the mechanisms for the next stage of improvement. Early deployments of visual intelligence platforms have reported measurable improvements in construction safety. In one Singapore project, the contractor recorded a 10-fold improvement in overall safety scores within months, attributed to more consistent detection of safety violations that often occurred between manual inspection rounds. The vision-based shift that is already underway What VLM-based AI agents represent is not just a better monitoring system. It is the first time the construction industry has been able to close the loop between visual data and safety action at scale, in real time, across every input source on the site simultaneously. Singapore’s construction sector has been watched for years. What it now has, for the first time, is AI that genuinely sees it. Twenty workers died on Singapore’s sites in 2024, when cameras were running. The data was there. The comprehension was not, and that is what changed. That is why industrial safety, on the sites where VLM-powered AI Agents are deployed, has not been the same since. — Editor’s note: e27 aims to foster thought leadership by publishing views from the community. You can also share your perspective by submitting an article, video, podcast, or infographic. The views expressed in this article are those of the author and do not necessarily reflect the official policy or position of e27. Join us on WhatsApp, Instagram, Facebook, X, and LinkedIn to stay connected. The post The vision-based shift: Transforming construction safety with AI appeared first on e27.

Author: Gary Ng

Source: e27