What really really surprised me most is how little of computer vision is actually about smart code. Frame differencing and background subtraction are pretty dumb algorithms on their own nothing like how a human eye and brain read a scene using memory and context. Levin’s whole point is that you make the physical world more predictable instead like backlighting a silhouette and retroreflective material. That is the technique doing the work that smarter software would otherwise have to do, which changes how I used to think about computer vision as something you solve purely in code.
The surveillance question is where this got more complex for me. Suicide Box is unsettling specifically because frame differencing was enough to passively record something as serious as someone jumping off a bridge. It is not sophisticated AI which somehow makes it feel more invasive, like tragedy got reduced to a pixel threshold. Sorting Daemon and Standards and Double Standards push that further since both use the same basic tracking logic as their actual material, sorting or following people without consent. What struck me is that the exact techniques Levin teaches as beginner friendly are structurally identical to what gets used for surveillance. The difference is not code sophistication but intent and whether people know its happening.
That made me think differently about using webcam input in Hydra (in my live coding class i took last semester). I have treated camera feeds as just another visual source to feed into feedback loops without thinking much about it. After this, a simple motion detector pointed at a crowd is doing the same basic operation as Sorting Daemon, so if I use webcam tracking live going forward I want people in the room to actually know they are part of the input instead of just grabbing frame differences because it looks wayyy cooler!!!!