Human vision is one of the most mysterious things about the human brain. We don’t completely understand how the brain takes information from light reflecting off objects and turns it into something we can understand. Humans can look at what is in front of them and distinguish between different things in split seconds. However, computer vision is different because it works with digital images as pixels and uses specific rules and algorithms to understand what it is seeing. Because of this, computer vision is more restricted in what it can distinguish and is more dependent on the conditions of the environment.
As mentioned in the article, there are multiple techniques that can be used to track things, and I think the most important thing to take into consideration is the contrast of colors, light, etc. If there is a noticeable difference in these values, then the computer will be able to tell things apart more easily. The second most important thing is to tailor the rules and “physics” to the purpose of the program. For example, the frame differencing technique detects movement by comparing each frame to the previous one and looking for differences in the pixels, so to detect these changes a high contrast would be better.
I think computer vision has a huge capacity as long as you have lots of cameras and processing power, which is good when an interactive work requires lots of detail. For example, in virtual reality, to be able to detect your body movements, you may need multiple angles. Also having a big capacity can make the outcome more accurate and sensitive to small movements. But it also has its cons, because to some people it may be a breach of privacy. Doesn’t it make you uncomfortable to know that a computer has all the data of your body movements? That’s one of the reasons I liked the use of silhouettes in some of the interactive art examples. I think it’s a better way to give the same interactive experience but in a more comfortable way.