Oh no. my 2nd paragraph was about the raw depth data. From what I see, the web-cam view doesn't seem to contribute to the depth data (firing up the streams and viewing them shows different perspectives and there isn't any way to query how far a pixel in the web-cam view is from the camera).
I remember reading somewhere that the depth sensor projects something (infra-red?) on a surface and then the transformation of the image is used to build the 3D model. So it might still need some specialized hardware (not necessarily MS hardware since Asus seems to have a similar device out in the market).
EDIT : NVM I just looked up the video of Andrew Ng's class where they discussed this. So it is possible.
I remember reading somewhere that the depth sensor projects something (infra-red?) on a surface and then the transformation of the image is used to build the 3D model. So it might still need some specialized hardware (not necessarily MS hardware since Asus seems to have a similar device out in the market).
EDIT : NVM I just looked up the video of Andrew Ng's class where they discussed this. So it is possible.