How can we see a three-dimensional object in a flat image? Glasses and stereoscopes offer technological solutions. Autostereograms, like optical illusions and trompe-l'œil, require no device or apparatus.
Our brains' ability to reconstruct 3D representations from two almost identical images, one presented to each eye, has been known since the second half of the 19th century. The stereograms of the time (consisting of paired images; see the article "3D television: the television of tomorrow?") were produced using cameras with two lenses spaced as far apart as our eyes (6 to 7 cm). A simple optical device—a stereoscope—then made it possible to view the paired photographs and recover the sensation of depth. This is still the technique used in our children's View-Masters, as it has been for more than half a century, and by our television sets when they play so-called 3D Blu-rays.
In all these cases, creating the illusion of depth requires an aid such as glasses or a stereoscope. The challenge of autostereograms is to recreate that illusion without any such aid. Admittedly, all of us can naturally see in three dimensions (barring a particular impairment). But learning to see the magic of autostereograms takes real practice and, for some of us—though sadly not all—pays off with surprising hidden effects. A range of tricks and methods can make the learning process easier, notably through convergence points or optical aids (3D viewers, prisms, anamorphoses or LEDs). The possibilities offered by these images are immense; it is surprising that they feature in so few applications. Why not imagine an environment in which fences, screens, openwork partitions and wallpaper incorporate this kind of illusion, creating a whole new form of playful 3D design?
A new form of playful 3D design
----------------------------
We naturally perceive depth and three-dimensional shape. Our eyes converge on the same point, focus on it and send two images to the brain for interpretation by our visual cortex. We need only choose: in the example opposite, we can look either at the water tower or at the fence. We cannot see both planes in sharp focus at once.
Both eyes therefore turn simultaneously towards the chosen object. The nearer it is, the more our eyes turn inwards: this is convergence. They then focus: the lens in each eye changes shape to produce a sharp image. The nearer the object, the more convex the lens becomes: this is accommodation. Our brain receives and compares two slightly different images.
The nearer the object, the more the images differ: this is binocular disparity. After interpreting the information it receives, our brain gives us a sense of depth.
A flat image that produces a 3D effect cannot conform to these physiological principles, because our eyes do not converge and accommodate on the same plane. "Seeing in 3D" therefore requires effort from the viewer. Some people will see the three-dimensional image easily, while others will find it more difficult. For still others, it will be impossible without assistance.
The challenge, then, is to make our eyes fix on a point outside the image without accommodating to the image itself. Binocular disparity requires either a pair of stereoscopic images or an autostereogram for the brain to analyse. Convergence prompts both eyes to fix on a point either in front of (AV) or behind (AR) the image, causing the brain to overlap the left- and right-hand patterns and fuse them into a single image. There is, in fact, a relationship between the eye–AV point–AR point distances and the image's pitch. Lastly, accommodation requires us to give our visual system a few seconds to eliminate the double image and bring it into sharp focus.
####
Constructing an autostereogram.
To create the illusion of depth, an autostereogram consists of a single image, but it is generally made by horizontally repeating a modified version of the same pattern. Variations in the spacing between corresponding points in the patterns enable the brain to detect the hidden 3D shape. One important quantity is therefore the autostereogram's pitch, the distance between two equivalent repeating patterns. In the kangaroo pattern opposite, the pitch is the width of one kangaroo. The hidden 3D shape—which readers can discover for themselves in what follows—is like a wave or a lemon squeezer, appearing hollow or raised depending on how it is viewed. Download the Kangaroo tiling and instructions.
A question of convergence
---------------------------
To see the stereoscopic image in 3D, our eyes must converge on the same point, either behind or in front of the image, so that our brain can fuse the two images into one. This convergence can be represented geometrically by the angle between two line segments, one starting at the left eye (Og) and the other at the right eye (Od), and meeting at a particular point of convergence (PdC), located either in front or behind.
To produce the 3D effect, these two line segments must necessarily coincide with the pitch p of the stereoscopic pair. When viewing behind the image, the left eye must look at the left-hand image and the right eye at the right-hand image. This is known as parallel viewing. When viewing in front of the image, the left eye looks at the right-hand image and the right eye at the left-hand image; this is cross-eyed viewing. These terms can be misleading, so convergent viewing in front (AV) and convergent viewing behind (AR) are often preferred. Switching viewing modes reverses the three-dimensional form: a "hollow" becomes a "bulge", and vice versa.
The return of Thales' theorem
------------------------------------------
Given that the stereoscopic image must be parallel to the plane of the eyes, Thales' theorem can be applied to the triangle formed by the points Og, PdC and Od. It then gives us a relationship between the following quantities:
\- The distance d between the viewer's eyes, which varies from about 6 to 7 cm depending on the individual's anatomy and age, with an average value of approximately 6.5 cm;
\- The line segment p, which defines the image's "mean pitch" and can be approximated by (pitchmax + pitchmin ) / 2, since the pitch varies across the image. Otherwise, we would simply have an image made up of regular patterns, with no 3D effect. The greater the variation in pitch, the greater the depth;
\- The distance a between one eye and the vertical plane of the image;
\- The distance b between the image and the point of convergence PdC;
\- The distance c between one eye and the point of convergence, with a physiological lower limit that varies from person to person and below which sharp focus is difficult (25 cm on average).
Above all, we want the 3D effect to appear quickly. It is therefore best to focus on the centre of the image. In that case, we can assume that the distances a, b and c "do not vary too much" from one eye to the other.
Now imagine two stereo views, left–right (G–D), brought closer together—the pitch p decreases—until they overlap (p = 0), then cross and continue moving apart, becoming a right–left pair (D–G). The change from positive to negative pitch—the reversal of images G and D—perfectly captures the reversal of the 3D form.
Mathematically, the relationships between these different quantities are expressed by the preceding diagram. In parallel viewing, Thales' theorem (see Tangente 167, 2015) gives us:
pd=bc=c−ac.
This relationship implies ad = cd – cp, or equivalently
ac=d−pd.
Let K3D denote this ratio, expressed as a percentage. We know that K3D > 100% because c > a. It follows that, for a given person, the ratio K3D is directly related to the image's mean pitch p.
The same reasoning can, of course, be applied to viewing behind the image, giving the relationship:
pd=bc=a−cc,
which now implies
ac=dpd.
Expressing K3D as a percentage makes it easy to estimate where the point of convergence lies relative to the eye–image distance. When viewing behind the image, c > a, and hence K3D > 100%. When viewing in front, since c < a, we have K3D < 100%. Thus, for a viewer whose eyes are 6.5 cm apart, looking at an image 40 cm away with a mean pitch of 3.25 cm, we calculate that, when viewing behind the image:
K3D (as a percentage) = d / (d – p) = 6.5 /(6.5 – 3.25) = 200%.
For an image 40 cm away, the viewer will have to make their eyes converge at a distance of:
c = a K3D = 40 × 2/100 = 80 cm.
For convergent viewing in front of the image, we find that K3D (as a percentage) equals d/(d + p), namely 6.5/(6.5 + 3.25), or 67%. Still assuming that the image is 40 cm away, the viewer must make their eyes converge at a distance c equal to a K3D, namely 40 × 67/100, or 26.6 cm.
The calculations are now complete. But this in no way means that we can locate an autostereogram's point of convergence without a calculator, a table or a 20 cm ruler. So let us return to our 3D image and temporarily forget about our second eye. We must determine how wide a screen q, placed at a distance c from one eye, should be in order to span exactly the angle subtended by the autostereogram's pitch p (located at a distance a from the eye). Thales' theorem applies to this configuration as well.
We must calculate the distance b at which to place the stick (of height h) so that it lines up exactly with the pyramid's shadow (the pyramid having height H).
Thales' theorem
For viewing behind the image, we can check that
pq=ac, and hence q=apc.
Given the preceding results, we obtain
q=d−pdp.
By symmetry, for cross-eyed viewing, we obtain
q=dpdp.
Consider an observer whose eyes are 6.5 cm apart and who is viewing an image with a repeat distance of 3.25 cm from 40 cm away. For parallel viewing, we find q = 6.5 cm × 3.25 / (6.5 − 3.25) = 6.5 cm.
For cross-eyed viewing, the required width is q = 6.5 cm × 3.25 / (6.5 + 3.25) = 2.2 cm. The width q (the projected repeat distance) is directly related to the image's repeat distance. This relationship can be represented by a universal "3D viewer" (a pair of intersecting hyperbolas), which lets you read q directly from the x-axis given an average repeat distance on the y-axis.
The illustration below shows both configurations: cross-eyed viewing (AV) and parallel viewing (AR).
How far can this type of image be taken? Parallel viewing is limited by how far apart the eyes are, so it cannot accommodate “large” images—more than one metre wide, say, or sixteen repeats at the maximum 6 cm interval. Cross-eyed viewing has no such limitation, but fewer people use it.
Incorporating the 3D viewer, or a physical marker of the convergence point, into the setup could open up new applications. One final tip for seeing the 3D effect more clearly: choose a bright environment for better contrast, and position the image to obtain even lighting, if possible without shadows or unwanted effects. Now it's your turn to practise!
Theory over—now for practice!
-------------------------------
The paradoxical brick wall is a paper model to fold using the downloadable instructions. This example shows how the brain can be fooled into perceiving the apparent three-dimensional form as its opposite.
Depending on how it is viewed, the paradoxical spring may appear either as a 3D figure or as an impossible figure.