Reading Reflection and Midterm progress report

Reading Reflection

The Minsky story at the start is funny, but I think it explains why computer vision is so hard to appreciate. Seeing feels effortless to us, so it seems like it should be simple, when really it’s one of the things we understand least about our own minds. Levin points out that text already comes as symbols, while video is just rectangles of pixels with nothing telling the computer what any of them are. What I kept thinking about is that human vision isn’t assumption free either. We fill in gaps, fall for optical illusions, see faces in random objects. The difference is that our assumptions are invisible to us, while a computer’s have to be written down by someone. In LimboTime, the program treats the limboer’s head as the topmost point of the middle blob of black pixels. The computer isn’t seeing a head there. A few workshop participants decided what a head is for one afternoon, and the computer can only ever find what they told it to look for.

Looked at that way, each technique in the reading is a different guess about what matters in a scene. Frame differencing guesses that what matters is whatever moves, which is why it works on the Tour de France and fails in an office waiting room. Background subtraction guesses it’s whatever wasn’t there before. Thresholding guesses it’s whatever is darker or lighter than everything else, and brightest pixel tracking is basically a max search over an array. None of this is hard to code. The harder part, and what Levin spends most of his time on, is making the world fit the guess with backlighting, infrared, polarizing filters, or retroreflective material (the same stuff 3M makes for safety uniforms). His conclusion says a scene’s “quality” is defined by the algorithm analyzing it, which means there’s no such thing as a good image in general, only an image that’s good for one particular piece of code. The idea I found most useful is at the end of section three, where he says to choose features that are easy to detect and also carry meaning. The mouth is an easy dark shape to find, and its roundness happens to track vowel sounds. Pupils reflect infrared and also show where someone is looking. Power of real design skill is embodied by the difference of what’s computable and what’s meaningful.

That same idea is also what makes surveillance work, and Suicide Box shows it most clearly. The camera never knows anyone is jumping. It picks up vertical motion through simple frame differencing, the same technique a beginner would use for a game, and the location of the camera turns that motion into a death. Levin mentions in the abstract that computer vision was mostly funded for military and law enforcement, and then his conclusion lists everything it can now track, including identities, facial expressions, gait and gaze direction. In 2006 that probably read as a list of creative possibilities. Reading it now, it sounds a lot like a description of the systems that watch us in airports and on streets.

Since the technology is identical, I think the line between interactive art and surveillance comes down to whether the person being tracked can see what the system sees and gets something back. Videoplace shows you your own silhouette, so you’re playing with the system instead of being processed by it. Sorting Daemon turns people on the street into color swatches, and most of them probably never find out. Cheese sits somewhere in between. The actresses knew they were being watched, and an alarm pushed them to make their smiles look more sincere to a machine. Every interactive piece with a camera does a softer version of this, getting you to move in ways the camera can read. Usually that’s the fun part, but Cheese shows how quickly responding to someone turns into making them perform. If I were to determine for what counts as a head or a smile, I would let people see what’s being tracked instead of hiding the camera and letting it feel like magic.

Midterm Report

I was thinking on notes of what is culturally important to me and values wise. Being from Pakistan, hospitality is a big part of the culture, and I believe I have embodied it in a nuance way. Whenever, my friends and I go out I want to pay for us. There is an element of looking after people that brings me a lot of joy, and I wanted to embody that in the game.

So, I will be making a game where the faster you approach the card reader you win. It will mainly be pressing a key to get to the card reader faster, while the other 2 players reach the machine.

I have prepared the 4 frames that will be needed for the game, and I have made a pixelated version of the frames. I wanted to have the first 3 frames just like appear and kind of build a storyline, and the fourth frame has the game in it.

Here is a wire-frame of the game:

Then, here are the 4 frames of the storyline. 


Parts I am worried about

When the user starts playing the game, I am concerned about how the hand would elongate. I will have to parse the image and maybe separate the hands from the arms like save it as a separate asset too. Then, as the user plays the hand extends towards the card machine, so i am concerned about the game mechanics too.

I will be doing something like following to move the separate assets:

progress = 0                        // 0 = start, 1 = touching the card machine

on key press:
    progress = progress + 0.05      // each press moves the card a bit closer

every frame:
    progress = progress - 0.002     // slides back if you stop pressing
    handX = startX + (machineX - startX) * progress
    handY = startY + (machineY - startY) * progress
    draw hand at (handX, handY)

    if progress >= 1: you pay
Progress so far

Leave a Reply