← colinfizgig · research

Essay · Revised 2026 · First published on Medium, April 2016

The Visual Language of Reality

Why AR and VR demand a visual grammar of their own, and how the way we already see could be its foundation.

Read the original 2016 text

An eye superimposed on a camera lens: the kino-eye of Dziga Vertov's Man with a Movie Camera.

About This Edition

I wrote this essay in 2016, while my research at Georgia Tech was focused on augmented reality and interactive narrative. In the ten years since, that research has turned toward the frame itself: how human perception composes images inside a rectangle, whether that structure is hard-wired or learned, and what it means for any medium that tries to direct where we look. Studying cinematography through perceptual psychology has changed my mind about some of what I argued here and sharpened the rest.

Rather than quietly rewrite the past, this edition keeps the original argument and marks where my thinking has moved. The biggest additions are a rethinking of what the frame actually does for film, and two new sections: The Frame Was Never Only an Edge, on what perception research suggests for immersive composition, and The Actors Have Arrived, on the AI characters I now build and direct. The 2016 text is still available exactly as it was first published.

Introduction

The history of human culture is told and explored through story. In the beginning the medium was the voice, as stories were told around a fire and passed on through oral tradition. Later the medium changed to text on stone or paper, and still the story was recited through an actor’s voice. With the advent of photography and cinema a new visual language began to be explored. The narrative techniques for staged live performance were slowly replaced with newer, more powerful techniques like editing, framing and compositing of imagery, enabling the artist to manipulate time and visual reality in order to tell the story in a whole new way.

The traditional stage methods of lighting and props were still used in films, but with the ability to slice a segment of film into parts and include totally different shots in sequences, the medium took an exponential leap in its ability to entertain and tell stories. These techniques have been further enhanced by the use of digital editing, compositing and virtual model making. We have come to the point where any fiction or reality can appear real: time travel, space exploration, death, re-birth; it is all within the realm of possibility when viewed through the eye of the cinematographer.

Now there is a new form of digital media brought about by the advent of digital graphics technology and cinematography called virtual reality. Virtual reality and its hybrid sibling augmented reality offer new, more immersive ways to tell stories. In many ways they should be the end of the evolution of story building, having come full circle to the point where the viewer can finally be the actor in the story if they choose to be. There is, however, a problem. The language of cinematography which enhanced film narrative to the level of art relies very heavily on the two-dimensional nature of film. The frame of the shot and the ability to cut from one view to another instantly without cutting through the viewer’s engaged immersion is key to the powerful effect good cinematography has on story. It took me years of research to understand why. Rudolf Arnheim argued in 1932 that film is only a partial illusion: we experience it as a real space and as a flat picture at the same time, and it is the picture half of that experience that lets us accept a jump from one place to another without feeling thrown. Perceptual research later found a second reason. Every time our eyes leap from one point to the next we are briefly blind, and a well-made cut rides on that same blindness. The frame and the cut are not arbitrary conventions; they are built on how we already see. Even theater was framed by the stage with the viewer’s perspective constrained to the lit view of the set and actors. The problem for virtual and augmented reality is that the medium surrounds the viewer and, in the case of augmented reality, overlaps the world around them, stripping away the very buffers that made the frame and the cut work. There is no flat picture to fall back on, and the viewer’s own eyes and head decide when the view changes.

How do you frame a shot when the story occurs in the world around you? How do you control a stage when the actors may be unknown individuals, on their way to work, walking between you and the focus of the story as it unfolds in the environment you both share? Is it possible to take a story intended for a specific place and allow it to occur anywhere? These are the questions that must be considered when thinking about narrative and visual storytelling as it applies to digital realities. There is also a newer possibility, one I could only gesture at in 2016: some of the actors may be agents, characters played by language models who remember, want things, and can take direction.

Some may argue this is a problem that is already solved. Games are virtual realities, and they have stories, sometimes multi-narrative stories applied to them. In 2016 most games, narrative or otherwise, were still played on a framed screen, and the few built for virtual reality leaned on the constraints below. In a discussion with most people who have made these types of digital realities you will often hear the word “tradeoff” used. The interaction might be constrained to an “on-rails” approach or the interactor is standing on a platform which just happens to fit the narrative of the story. Some more “exploratory” examples might constrain the user to a single room with one character and some props for exploration. There is a reason for the limitations applied to most virtual reality experiences. It is very hard to create an immersive, compelling narrative when the view is controlled by the viewer. Imagine a movie which showed the director’s view behind the camera with no ability to cut from one shot to another and you begin to see the problem. A decade later, room-scale headsets, hand tracking and mixed-reality passthrough have loosened many of those constraints, but the underlying tradeoff remains: the more control the viewer has over the view, the harder it is to direct them.

By exploring the theory and evolution of visual storytelling solutions for theater and film it may be possible to develop a new visual language which allows the new medium to evolve and enables even more powerful forms of narrative. Since the advent of story, narrative and location have been locked together. With the advent of augmented and virtual reality it should be possible, given the right visual tool set, to allow any narrative to be retold specifically related to any location. Obviously, this magnifies the problems of traditional cinematography when applied to augmented reality, since the stage or set changes depending on the viewer’s location. For that matter, who is the viewer in the narrative? Is she the main character, a third-person viewer like a traditional film, or an extra in the movie? This further complicates the problem. However, given the powerful techniques of modern film and digital media, these problems are not insurmountable. They simply require a new way of thinking about visual storytelling.

Narrative and Performance Studies

When film was first shown in cinemas the event was very similar to the production of a theatrical play. There was a stage and a screen on which the film was projected. And so the notion of how films could tell stories took on the methods of live theater. Some of the masterpieces of the time employ staged sets very much like a theatrical production would.

A chorus line loading a giant cannon on a painted stage set, from Le Voyage dans la Lune.
Le Voyage dans la Lune (1902) illustrates the obvious theatrical staging influences in early cinema.

Eventually film evolved its own language in cinematography and montage. Now it was theater’s turn to borrow from cinema. The “epic theatre” of Bertolt Brecht with its Verfremdungseffekt (translated as “defamiliarization effect”) uses several concepts such as a montage technique of fragmentation, interruptions in action, breaking the fourth wall and contrast and contradiction to allow the audience to feel detached from the play. The goal of these techniques was to promote rational self-reflection in the viewer. [‘Understanding Brecht’, Walter Benjamin. 1983.]

These types of effects might also be used within digital realities to allow the viewer to focus on the intended focus of the narrative. Elements within a scene could be slightly different or paused in action compared to the world around them in order to guide the user to the next sequence of events in a narrative. The thing that interrupted the viewer to create self-reflection for Brecht might serve well to grab their attention within an environment with no other frame of reference.

Perhaps one of the best visual examples of these techniques, as they might be applied to a virtual world, is embodied in a scene from The Matrix. The scene involves the teacher Morpheus taking the student Neo into the Matrix to explain how the system works. On this tour of the Matrix, Morpheus moves through a very crowded street when a woman in a red dress, the only thing red on the street, walks by and Neo is distracted by her. Then Morpheus stops the action of the scene allowing Neo to adjust his focus to the lady in question, only to realize with a shock that she is an enemy [The Matrix, the Wachowskis, 1999]. I originally described this scene as Brechtian, but I now think it is doing nearly the opposite. The red dress works through salience: she is the only saturated color on a street of dark suits, and color contrast is one of the strongest pulls on the eye that eye-tracking studies of film have found. Freezing the street then strips away every competing motion so that Neo, and we, can look again. Brecht’s estrangement breaks the spell to make an audience aware of itself. An immersive director needs both tools, but for opposite jobs: salience steers attention inside the story, and estrangement steps outside of it.

A woman in a red dress walks through a crowd of people in dark suits, from The Matrix.
The woman in red stands out from the crowd in The Matrix.

The concept of defamiliarization can be explained as taking something common or taken for granted and presenting it in an unfamiliar or strange way. The Russian literary critic Viktor Shklovsky declared this to be the essence of all art [ The Theory of Prose, translated from “Art as Device”, Shklovsky 1991 ]. The concept can serve well in the artifice of visual narrative as well, as witnessed in The Matrix.

If the red dress scene from The Matrix was a virtual reality, and you were in it, who would you be from a narrative point of view? You could answer Neo; it is the obvious choice since he does not talk and is following Morpheus, who is giving the speech. You obviously are not the lady in red since you notice her. There is another choice, though: you could fill the role of the observer, which is usually the role given to us in film and theater. Virtual reality transports us inside the fourth wall to the confines of the boundary as Janet Murray describes it in her book Hamlet on the Holodeck. The problem with being a third-person viewer in a virtual reality is the lack of interaction. You are like the viewer who has wandered up on stage, immersed, but somehow in the way unless the author writes the narrative with you in mind. Perceptual research gives this a physical dimension. Film almost never brings anything within a meter and a half of the camera; when it does, in the zone psychologists call personal space, we suddenly become aware of ourselves as viewers. In a headset, characters step into that zone constantly, so every choice about how close they come is also a choice about who the viewer is in the story. You could take on the role of the extra or supporting actor, but is this an interesting role that takes advantage of all that virtual reality affords? With agents playing the other parts, the role becomes a choice rather than a problem. In The Ships of Thespis, the character engine I have been building, a visitor can step into a character’s part, and the rest of the cast meet them as that character, with that face, that history, and whatever they already feel about them.

In a 2015 interview Ed Catmull, co-founder of Pixar and then president of Pixar and Walt Disney Animation Studios, warned that virtual reality technology is not storytelling. He stated that while virtual reality may have interesting contributions for gaming, he believed it was inappropriate for storytelling. He said, “Linear narrative is an artfully-directed telling of a story, where the lighting and sound is all for a very clear purpose. You’re not just wandering around in the world.” [ Interview with the Guardian. Stuart Dredge] This gets at the crux of his view, which has some merit but fails to consider the history of storytelling. He is comparing the art of cinematography to a new medium which has not discovered its own grammar. Virtual reality may try to mimic films and cinematography as a narrative medium, but eventually it will discover its own narrative tool set. In addition, what Catmull says is true: an artful story is well crafted and has a very specific designed and linear narrative. The concept of roaming through the story world seems to trample the concept of authorial intent, but does it really? In a poorly designed virtual reality it would, but what universal law of narrative prevents the viewer’s ability to freely explore a story world and the author’s intent for how the story unfolds? One “law” would seem to be pacing. If in the middle of a car chase scene a user stops the car and opens the door to look at a dead animal on the side of the road, the author’s desire to create a tense and heart-stopping action scene falls dead itself. In 2016 I called the author’s control of the viewer an illusion, and pointed to people who leave in the middle of a film to refill their popcorn. I no longer think that holds up. When people watch a film, eye-tracking shows their gaze converging on nearly the same point at the same moment, far more than when they look at a still image, and over the decades filmmakers have shifted their shot scales toward the ones that hold that convergence tightest. The author’s control of attention is real, and it can be measured. Where viewers differ is in what they prefer and what they make of what they see. So the challenge for virtual reality is not to abandon authorial intent but to rebuild that control with tools that work when the viewer holds the camera, and to leave the viewer’s freedom where it has always lived: in interpretation. Discovering an author’s intent by exploring a scene from multiple viewpoints can still be a gift to both the author and the viewer. Agents also change where the author’s hand sits. In a story played by characters who can act on their own, the author no longer writes every line. They author the people: what each character knows and must not reveal, what they want, what moves them, and the order in which their story comes out. And they direct, with a director that decides beat by beat who should react and what each is after. The viewer who stops the car to look at the dead animal no longer breaks the scene. The characters notice, react in character, and are steered back toward the chase. It is true that film’s language seems optimized and extremely specialized for telling stories, but to imagine that the language of film existed from the advent of photography is a fallacy. It had to be discovered over time. In fact, the language of cinematography is still evolving, and I suspect the people at Pixar would be the first to say so.

Consider your viewpoint if the red dress scene was an augmented reality. In other words, what if the red dress scene occurred in the real world around you? How is narrative viewpoint different in VR and AR? There is a subtle but powerful difference between a completely virtual environment and one which overlaps the real world. The real world has chaos, real objects to composite virtual objects against, real sensory experiences and many unexpected events. There are a million little things that create an impenetrable boundary wall within augmented reality narrative. This can be problematic for narrative but it provides affordances as well. One important affordance is that the scene does not have to rely completely on computer graphics for its environment; it can use the environment around it. The scene would not need hundreds of CG avatars and AI agents (programmed agents, not Agent Smith) to animate them. We could take the role of any character in the scene as long as we played the part. There are real perceptual issues with forcing alternative viewpoints upon a viewer in virtual reality which can cause physical discomfort and loss of immersion. These issues do not occur in augmented reality since most of the view is reality as the user’s body experiences it. Augmented reality narrative appears to dismantle the concept of an “artfully-directed” story, but does it really? Is an artfully directed story completely locked to the author’s intent to the point it can only occur in one time and place? This is the crux of the discussion for a visual language for digital realities. Many stories are intended for a specific place and time, but only because those stories were limited by their author’s intent and the medium for which they are created. Consider a dinosaur in a book, walking along talking to another character. A film director could add more to the scene in order to fill in details or require extra lighting if the event took place at night. The question is: does the extra lighting affect the author’s intent? What affordances would drive the virtual reality director to create some other twist for the scene allowing them to artfully direct the story for their own medium?

A line drawing of Gertie the Dinosaur beside Arlo from Pixar's The Good Dinosaur.
Gertie the Dinosaur (1914) and Pixar’s The Good Dinosaur (2015)

Film Theory

Is the “Subject” of a narrative an immutable object which cannot be represented in multiple ways? For that matter, what is the “Subject”? In “Contemporary Film Studies and the Vicissitudes of Grand Theory,” David Bordwell criticizes the way both subject-position theory and culturalism collapse the subject into the individual viewer [Post-Theory: Reconstructing Film Studies, eds. David Bordwell and Noël Carroll, 1996]. For my purposes the subject can be both: the individual within the story, and the knowledge the story carries. Using the dinosaur example from earlier, imagine moving that story into a modern metropolis. If the author’s intent was to show friendship between two very different species, the dinosaur might befriend a businessman rather than a caveman, and a more creative author might find a parallel between hunting for work every day and hunting for food. Changing the place and time changes the mechanical structure of the narrative, but does it change the ideas? This is no longer only a thought experiment. The characters in The Ships of Thespis carry what I call reality variables: change one, and Victor Frankenstein lives in a world where Justine was acquitted, and has always lived there. Whether the idea survives the change can now be tested scene by scene. Story, like all of reality, has parts that are essential to it and parts that are relative and changeable.

In 2016 I framed this as a debate between schools of film theory. My research has since turned it into the oldest question in the study of composition: which parts of how we see are given by perception, and which are learned from culture? The answer I find most useful comes by way of Henri Bergson. The structure perception finds in a frame is real and present in every viewer; the visual habits we build from a lifetime of images are just as real, but accumulated over time. Film’s grammar rests on both. The perceptual substrate is shared, and the vocabulary built on top of it is learned. An immersive grammar will be built the same way, which means its designers need to know which of film’s conventions are perception and which are only habit.

New Media Theory

A grid of black-and-white film frames laid out like a database of shots.

In his book The Language of New Media, Lev Manovich discusses the essential elements of digital media as being databases and algorithms. He compares this to early works of film such as Dziga Vertov’s Man with a Movie Camera in which Vertov cut footage and laid it out in large grid-like structures in order to plan and edit his media into a new narrative. In addition, Vertov used elements of montage and “kino eye” to alter the image of the original footage, creating a visual language which allowed the seemingly random database of shots to tell a story. Manovich points out that the linear structure of narrative and the non-linear structure of databases would seem contradictory, but that in effect linear narrative is really just one path through the database of a narrative’s fabula (everything within the narrative world). To illustrate his points Manovich enlists Ferdinand de Saussure’s concept of syntagm and paradigm as they were further expanded by Roland Barthes [Elements of Semiology, Roland Barthes 1964]. Manovich describes the linear narrative as the syntagm, the linear path through the world, of the author’s paradigm, all the imagined paths the author might have described [The Language of New Media, Lev Manovich 2001]. This, better than any other concept, can describe the viewer’s choices in a digitally constructed reality. The virtual reality must therefore be the author’s fabula or paradigm while the path of the user becomes the narrative or syntagm explored by the interaction of the viewer and the author’s skill in guiding the viewer on the path she desires for the main character in her plots. Whether the viewer follows the author’s direction or trots off to explore some other element of the story becomes irrelevant. The truth remains: a narrative exists, a more interactive narrative world exists and the viewer may read into it what they will. This is not a new concept. It has almost always been the case that the reader/viewer was the driver and the author was simply a helpful guide giving them a map allowing them to explore the world. Some authors are more capable than others, which gives the illusion that they direct the viewer completely. Some readers are less inclined to diverge from the path of the story created by the author, but still they have their own interpretation of the story to follow. New media simply brings about a swap in the status quo: the author creates a paradigm and the reader picks the syntagmatic path of their choosing. Where the database story world becomes the main tool for the author creating paradigmatic rules for her fantasy world, the algorithmic deciphering of rules for that world carried out by the viewer becomes the artfully directed path of the viewer.

Looking back, this is the part of the essay that has most shaped my research. If a narrative is one path through a database, the database can be indexed by more than its content. My current work trains a vision-language model to annotate film frames by where their subjects sit within a composition grid derived from each frame’s aspect ratio, turning a film into a searchable record of compositional choices. Shots could then be found and sequenced by how they are composed as well as by what they show, and the path a viewer takes through a story world could be measured rather than only described. My proposal for a “synthetic cinema,” assembled from found footage by a script, machine learning and composition grids, is a first step in that direction.

Agents make the database perform. In The Ships of Thespis each character is a package of psyche, memory, canon and world: a paradigm in Manovich’s sense, every path the author has made possible. Each encounter is one syntagm through it. When several characters share a scene, a final editing pass may trim and interleave their lines, but only with words they actually spoke; it may cut and it may reorder, but it may never rewrite or add. That is montage performed on a database of performance, about as close to Vertov’s method as a live medium can come.

Augmented and Virtual Reality as a Narrative Medium

What are the narrative affordances associated with the database and algorithmic nature of digital media as well as virtual and augmented reality? In the beginning of the 20th century, filmmakers began to create their own visual language and techniques which enabled them to change the nature of the stories they could tell. For example, a linear series of events could be cut and reordered to create a completely different perception of the story and event order. Suppose that everything about a specific time and place could be recorded as a series of events and media elements. This would enable the author to rearrange that database into any order and output various stories based on each new sequence of events. This is the advantage and essence of film as it relates to narrative. However, film does have its limits, in that it must be processed and can only be shot from certain angles. The framing of a film cuts off potential views of the environmental elements of a story, thus a master of cinematography becomes an artist known for their ability to create a masterpiece with every frame. That is the dogma of film and cinema: there can be no other view but the view framed and chosen by the cinematographer. It is tempting to call this a lie, given the hours of footage left on the cutting-room floor, but Rudolf Arnheim offered a better reading back in 1932: a medium’s limits are its tools. The frame, the single viewpoint and the cut are not failures to capture reality; they are what make film an art. Obviously, the immersive nature of virtual reality allows the viewer to become the cinematographer, which with some guidance and visual trickery could be directed by the professional using elements of interface and visual language to allow the viewer to discover the best viewpoint for themselves. The real question is not whether the cinematographer’s choices were a kind of dictatorship, but what will do the frame’s work once the frame is gone. With VR/AR the power is handed back to the individual, with the director or author serving the role of masterful guide in their story world.

VR/AR gives the individual even more narrative control than they already possessed while enabling the cinematographer the opportunity to free themselves of the frame and attempt to see the world as a more artful reality. This, of course, comes with some very weighty pitfalls and problems for the medium, which can literally make viewers sick given its pervasive inputs to their brain. For example, a cinematographer is free to establish shots at a distance and fly the character to the close-up perspective they intend for a scene. They may even choose a locked shot hovering thirty stories off the ground looking through a window. These artful methods are steeped in signs and symbols for the director and the audience. However, these same techniques become problematic within digital realities because they may break the viewer’s sense of immersion and physical engagement with the scene. Arnheim’s partial illusion explains why. On a screen, the flatness of the picture absorbs a sudden change of viewpoint; in a headset there is no picture left to absorb it, so the body takes the jump literally. The eagle-eyed view in film becomes the unintended vertigo-inducing effect in reality. There seems to be no way to alter this without the use of props like platforms or avatar bodies which allow flight. These workarounds begin to weigh down the potential world rules with mechanical “filters” which tie the user’s imagination down and limit the author’s designs. This, like the framed view of film, becomes the advantage and disadvantage of digital realities, at least until we cross over into the Matrix and believe like Neo that we can jump across the gap and that flight is natural. Language is bound by rules and limitations; that is what allows understanding and a sense of wonder when the rules are broken and the limitations are forgotten. It is the artful manipulation of those limits which sets the cinematographer apart from the home videographer. It may be for these reasons that the visual language of digital realities will rely more on effects like Vertov’s “kino eye” and alternative forms of view manipulation like those used by magicians on stage and modern theater in order to artfully guide a user through a scene.

How will a viewer be directed from an establishing shot to the desired focal point of a scene in virtual and augmented reality? Obviously in virtual reality there is the “on-rails” approach and many VR narratives already employ this technique. This is one approach but is severely limiting in a medium which should seek to remove limits on the viewer. In addition, this approach will not work for augmented reality since the user cannot be constrained in the real world. How then would you lead a viewer to the focal point for a scene in augmented reality? If this particular problem could be solved then the framework for a digital reality visual language can be set. Solutions include directing the viewer with a visual overlay or path, not unlike the subtle dodging and burning used in film to guide the viewer’s eye to the desired focal point of the frame. There is also the possibility for 3D sound which could enable a prop or character within the scene to call to the viewer in order to get their attention. In addition, the compositing of the scene could be controlled in such a way so that the element of focus could be the only element of the scene in color and the rest of the scene could be black and white until the viewer moved to the perimeter of the shot. Now the methods for exploration start to look more artful, similar to the fade from black of film, the iris close on the main character at the end of a scene used in film or the spotlight used to highlight the speaking characters on a stage. In 2016 I suggested that the frame of the 2D film becomes the perimeter of the 3D scene. I now think the frame’s real power was never its edge. A bounded frame generates an invisible structure inside it, what Arnheim called its structural skeleton: a center, axes, and positions where a subject feels at rest or in tension. Stephen Palmer’s experiments confirmed that this skeleton can be measured, and found something even more useful for immersive media. When a frame is divided by an interior border, each part generates a skeleton of its own, strong enough to override the whole. Doorways, windows, arches and the gaps between buildings are frames within frames. An immersive director cannot frame the world, but can stage the world’s own frames, and the viewer’s perception will organize what appears inside them. Oskar Schlemmer understood something like this at the Bauhaus in the 1920s, when he strung taut wires across his stage to make the invisible geometry of the space visible.

What about the other affordances of database for digital realities? Film offers the affordance of recording ‘framed’ time. While reality has space, the time associated with it only moves forward. The history of the location in reality is buried or forgotten depending upon the viewer’s knowledge. Film’s linear cause and effect can be viewed forward and backward. Could there be a rewind for reality? Yes, and perhaps much more. Human culture is beginning to establish databases of time and space for large chunks of reality. When I first wrote this, Google Earth, Street View and YouTube, fed by cameras like the GoPro, were the best examples of our desire to record and catalog reality in as much detail as possible. Since then, geospatial anchors can pin virtual content to a particular street corner, photorealistic 3D maps cover whole cities, and Gaussian splatting can capture a real place with little more than a phone. One highly underused affordance of virtual reality and augmented reality in particular is the ability to tap into the ever-growing database of reality and use it to construct narratives. Let us imagine that the digital media revolution began when the film camera was created and that we have access to all the media created since that point in the form of a spatial database associated to a map of the world, similar to Google Street View. The virtual auteur could take an algorithm for artificial intelligence capable of parsing that record of space and time and use it to construct a narrative for an individual based on their current location. Obviously there is a lot of work that goes into the creation of this type of tool, but the building blocks are already there and the potential for narrative and visual language is enormous. A decade later, generative models that can read and produce images and video have made that sentence far less speculative. This is one of the reasons I believe that the concept of virtual reality and augmented reality will eventually merge into one medium which is capable of both types of outputs. That merger is now well underway; today’s headsets move between virtual reality and passthrough mixed reality within a single experience. There are too many compelling uses for the database-driven augmented reality narrative for it to play second fiddle to the less useful but more robust power of virtual reality. You could compare it to the difference between film in a theater versus the combined viewing of Netflix and YouTube cut into viewing segments based on sequences of actions and effects, then re-sequenced into any combinative story imaginable.

The affordance of massive re-combinative media is extremely useful from a content perspective, but what does it mean for the proposed visual language of digital realities? These types of inputs become the foundation for rules that drive the compositing effects of digital realities. Modern film effects creation relies heavily upon 3D data. The same goes for digital reality compositing. In order to insert objects into a world scene, the compositor needs a 3D frame of reference. Taking geospatial positioning and databases of 3D data such as buildings from Google Earth, it is possible to create the same types of compositing which occurs in film in augmented reality in real time. In addition, using positioning, mapping, artificial intelligence and computer vision tools, it is possible to find and track locations within the real world which map well with a director’s intended location for a scene. For example, imagine that a scene requires the corner of a building and a side street in order to work well visually, as imagined by some AR director, human or AI. The nearest best match for that requirement can be tracked and located through algorithms and databases, enabling other elements of the “visual language” to then guide the user to that location in order to lay the scene out around them and continue on with the story. My research suggests how that match could be judged: a model trained to recognize compositional structure could rank candidate street corners by how well their own architecture frames the scene. Alternatively, the scene could be constructed to fit the surrounding environment or constructed completely in the case of virtual reality.

The Frame Was Never Only an Edge

If I were writing this essay today, this is where it would lead. The question running through it is how to direct a viewer who holds the camera. Ten years of studying cinematic composition suggest an answer: immersive media should borrow film’s mechanisms rather than its conventions.

Film’s grammar works because it is built on perception. John Huston observed in 1965 that the techniques filmmakers had learned so laboriously “were already part of the physiological and psychological experience of man” long before film existed. A cut stands in for the leap of the eye. A composition works with a structure that perception already finds inside a bounded rectangle. An immersive grammar should start in the same place. A turn of the head is the viewer’s own cut, and researchers have already used the brief blindness of eye movements to rotate a virtual world without the viewer noticing, steering people as they walk through VR [Towards Virtual Reality Infinite Walking: Dynamic Saccadic Redirection, Sun et al. 2018]. Architecture can do the work of the frame. And the viewer’s distance from the action, whether across a plaza, across a room or within arm’s reach, can do the work of shot scale.

The frame’s structure also changes with its shape. In my doctoral research I have been formalizing this as an Adaptive Grid: a set of compositional positions and lines derived directly from a frame’s aspect ratio, so the same structural logic applies whether the frame is the Academy’s 1.37 or anamorphic 2.35. James Cutting’s analysis of some fourteen thousand frames spanning seventy-five years of cinema found that filmmakers place characters in the same proportional positions whatever the aspect ratio, which suggests that the structure lives in the proportions of the frame rather than in any fixed grid. If that holds, then a headset’s field of view, a doorway or a window each carries a skeleton of its own, and an immersive director can compose for it.

Film also gives us a way to test whether any of this works. Eye-tracking shows that a well-made film pulls its audience’s gaze together, and that synchrony can be measured in a headset as readily as in a theater. A grammar for immersive media does not have to be argued into existence; it can be tested. Nor should it aim for perfect control. The composition research I work on predicts that the most compelling images depart from the underlying structure in measured ways rather than conforming to it exactly. The goal is a scaffold with room to wander, not a set of rails.

This is the framework, if you will pardon the pun, that I believe immersive media can build its own cinematic language on: perception as the foundation, the frame’s skeleton carried into space, and the viewer’s gaze as the measure of whether it works.

The Actors Have Arrived

In 2016 I mentioned AI agents almost in passing, as a way to fill an augmented reality scene without hundreds of animated avatars. Ten years later they have become the most practical part of this argument. I have been building a company of them: The Ships of Thespis, characters rebuilt from books part by part. Each is authored in layers: a psyche that says how they behave under threat, what moves them and what lines they will not cross; a canon of what they know and cannot know; memories of their world that surface when they become relevant; and an inner life of mood, intention and patience that carries from one beat to the next. A director reads each moment and gives every character a private note to play. A narrator describes the room but never anyone’s mind. An early version runs in the browser.

This matters for the grammar I have been describing, because agents are performers who can take direction. Everything this essay says about guiding a viewer, from salience to frames within frames to the head turn as a cut, assumes someone has arranged the scene. With agents, the scene can rearrange itself around the viewer, in character, while the author’s constraints still hold. The director becomes a cinematographer of behavior: deciding who speaks, for how long, who interrupts whom, and when a silence should be left alone.

The remaining barrier is not technical. It is whether audiences will accept a performance generated by a machine as legitimate storytelling. Film faced the same charge. Rudolf Arnheim wrote Film as Art in 1932 to answer critics who held that a mechanical recording could never be art, and his answer was that the art lives in how the medium’s constraints are shaped. Ed Catmull’s “not storytelling” is the same objection in a new form. Jay David Bolter and Richard Grusin called the process remediation: new media earn their place by refashioning the ones before them, as film refashioned theater, and as these characters now refashion both. I believe acceptance will come the way it came for film, not by hiding the machine but by making the authorship visible. In The Ships of Thespis the director’s notes and each character’s inner state can be opened mid-scene, and the samples of a character’s voice drawn from their book must appear in it word for word, inside their own dialogue.

The next step moves the author’s seat itself. In a workshop built on the same engine, a character modeled on Shakespeare serves as dramaturg for a new play. He vets the pitch, lays out the acts, writes the scene briefs and gives notes on each draft, but he never holds the pen. A separate drafting model writes each scene to his brief. Gates check it for borrowed lines, anachronism, meter and continuity. Then a simulated period audience watches the result and reacts only in behavior: where they laughed, where they grew restless, whether they applauded. The play is mended from that the way plays of the period were, between performances. Two plays, a tragedy and a comedy, have come through it so far. They are not Shakespeare, and every page says so, but they point to a different answer to the question of authorship: a company in which a person sets the task, the agents play their parts, and the whole process stays on the record.

It also gives new weight to the words from Roland Barthes that close this essay. These characters are, quite literally, a tissue of citations.

Conclusion

It should be noted that this essay makes many assumptions and chooses a very specific path to explore the perspective of a narrative within augmented reality. The reason for this choice is to push the questions and the answers for more complicated scenarios. The assumption being that if it is possible in augmented reality, it is easily possible in virtual reality, since the author has much more control over the story world and can craft it as specifically as a film director chooses sets and shots. The goal for both mediums is to consider the language of film and the language of new media in order to consider what hybrid child they might make.

There should be no question as to whether virtual reality and augmented reality are a powerful narrative medium. There are valid concerns and considerations for the medium’s approach to storytelling. These are no different than the initial learning steps which other mediums like film and video games have had to experience before they discovered their own grammar. The potential for the technologies discussed within this essay, specifically AI and its ability to drive this new visual language, has far-reaching implications for our culture as a whole. As it has often occurred in the past, new mediums come along and alter the way a culture consumes art and entertainment, which alters the way culture creates art and entertainment, which then goes on to create entirely new forms of media and new mediums to explore. It is easy from this viewpoint to imagine a future where a person can walk through a city experiencing a documentary about the buried history of the streets they walk on, not simply a linear narrative hard-coded to tell the same story every time, but a narrative made different with every step and turn the viewer takes. A walk that starts as a trip down memory lane may end up as a dinner murder mystery or a walk through the park and a Shakespearean play like A Midsummer Night’s Dream. We would not see them the way we experience these separate events today, but as one continuous flow of reality driven by the user and all the content at their disposal.

Writing about the hypothetical death of the author, Roland Barthes wrote, “We know the text does not consist of a line of words, releasing a single “theological” meaning (the message of the Author-God), but is a space of many dimensions, in which are wedded and contested various kinds of writing, no one of which is original: the text is a tissue of citations, resulting from the thousand sources of culture.” [ The Death of the Author, Roland Barthes 1967]. The truth of these words can be seen across mediums and time as books become plays, which then become movies transposed upon the modern world, which then become musicals or games, only to start the cycle anew. Creativity and imagination, along with a standard vocabulary for visual storytelling appropriate to the medium and grounded in how we actually see, are the only limiting factors preventing artful narrative within the realm of digital realities.