The project

A module for content

The Breaking Bad example

The specific point of reference in our approach to TV series was Breaking Bad, a show by Vince Gilligan that has become famous all over the world for being one of the most innovative series in terms of content and parallel plots. Thanks to these same features, it has been the perfect starting point to model a knowledge graph able to describe the content of any cultural property depicting a plot and some characters.
Analyzing the TV series different features, we found in the classification and description of characters the most interesting and richest field of research.

Competency Questions

In order to develop the new content module - the eighth within ArCo’s network - we adapted to our purposes the so-called method eXtreme Design, the same used for ArCo’s network.
Initially, we formulated a great variety of questions in natural language that helped us in identifying the most different aspects of Breaking Bad that seemed worthy of interest to us. A few examples are:

  • How many non-white characters are employed in illegal occupations?
  • Which female characters have both an occupation and a child?
  • Which are the characters that have been killed by other characters?
  • Who is the protagonist? Who are the fictional characters that are not the protagonist?

We then identified in each of them the logical components necessary to correctly formulating our competency questions, highlighting the events, classes, individuals (named or anonymous ones), relations and logical implications.
In particular, we ended up in defining the ‘Origin’ and ‘Profession’ classes (and their subclasses).
A more exhaustive representation of the steps occurred in this process can be consulted in the dedicated "Documentation" section or here.

Conceptual Model

A major and fundamental step in the construction of our content module was the realization of our conceptual model. the realization of a series of competency question was fundamental for us to highlight and underline the major aspects of interest in our cultural property and, even before formalizing them in the abstraction of a dedicated ontology, we carried on an additional step in the classification of the features we would have needed to deal with. Indeed, in order to understand in depth the relationships and interactions between the different narratological and sociological characteristics of our characters, we developed a conceptual model in natural language so that we could give a clearer visual representation of them and underline their features.
This process went through several steps of redefinition and curation, we were really focused on trying to have concretely the best match for our ideal conceptual model. Following it's presented a close up version of our conceptual model focused on only the major and principal classes; the full versione of it can be consulted in the dedicated "Documentation" section or here.


The Content Module

With all classes, subclasses and properties clearly linked, we were finally ready to translate the concepts expressed in natural language into a formal one. The whole process of building our modular ontology was carried on through the ontology editor and knowledge management system Protégé, through which we handled the task, where it was needed, of the placement on the one hand of some class restrictions in order to better define their application to real entities, and on the other one on property characteristics able to define the type of relationship they represent.
Our ontology counts 60 Classes, 51 Object properties and 12 Data properties, all specifically thought for the modeling of a cultural property content and particularly for the description of the characters. To see some example on how Protegè manage and represent our model costitutive element click here.

Knowledge Graph

The final step in this major second section in the development of our project was to convert the whole structure we have realized thanks to Protegè into a more clear and formalized way. For this purpose, we have created the LODE visualization of our module, which can be seen in its entirety here. From that visualization, we have extracted the knowledge graph automatically created using the software WebVOWL that allows us to have an immediate look at a glance at the nice structure of the module we were able to realize. 
This final action allows us to go ahead in the enrichment of our documentation, bring our proposal closer to a correct formalization, and make it an applicable module.

SPARQL Queries

Consistency checking

Once we built our ontology, we used it to create a dataset populated with named individuals derived from the chosen cultural property. This passage was fundamental in order to check the consistency of the model we had created against its application to real data.
Starting from the competency questions we listed at the beginning of the research, we translated them in a formal language to realize a set of SPARQL queries to be run on the dataset. They are built in different ways and on different difficulty levels, aiming at plumbing some of the relevant aspects present in our model.

Query 1:

How many characters not-white are employed in illegal professions?

Query 2:

Which female characters have both an occupation and a child?

Query 3:

Which characters have professions somewhat related to drug?

Query 4:

Who are the fictional characters that are not protagonist?

Query 5:

Who are the antiheroes that have work relationship the villain?

Query 6:

Who are the opposing characters with scope not universal good?

Query 7:

Who are the characters that has been killed by protagonist?

Query 8:

Which are antihero's sociological characteristics?

Other sources

Reliability checking

We decided to conduct additional SPARQL queries as similar as possible to ours on much larger and more authoritative databases such as the one consisting of data from the Wikidata Ontology Project dedicated to Breaking Bad and its characters.
We noticed that the point of contact were only partial: it was possible to make comparisons only in case of queries referring to the profession of the characters, their sex or, in some cases, the ethnic group they belong to; the information regarding the narratological role of the characters is almost absent in Wikidata.
We believe that this distance is a symptom of how much research on character descriptions is still at an early stage and, at the same time, worthy of being conducted. For these reasons we think that our module is original and worthy for potential expansion.

Wikidata Query 1:

All Breaking Bad non-American characters who are drug dealers.

Wikidata Query 2:

All female characters in Breaking Bad who have an occupation and a child.

Wikidata Query 3:

All BrBa characters who have an occupation related to drugs.

Wikidata Query 4:

All Breaking Bad character who have narrative role not protagonist.

Wikidata Query 5:

All BrBa characters who have work relationship with other characters.

Wikidata Query 6:

All Breaking Bad characters who are against some character (and who).

Wikidata Query 7:

All BrBa characters who are killed by another character, and by who.

Wikidata Query 8:

All Breaking Bad protagonist sociological characteristics.

ArCo’s spin-off (a proposal)

We imagined our Content Module as the eighth one in the ArCo’s network, a self-standing, independently working section of the whole project. We are aware that there are not many points of contact with ArCo itself: however, we do not feel we have failed, because we understood from the beginning many of the difficulties we encountered.
The meaning behind “ArCo’s spin-off” is, in fact, exactly this: an extension proposal derived from the intention to explore the possibility, for ArCo, of opening up to the description of the content.