Thursday, May 23, 2013

Much Ado About Nothing: Blank Nodes in RDF

Here's a secret - you want to understand a data format? Learn its query language. I've worked heavily with XQuery for several years now, but only fairly recently (three years now) did I start working with SPARQL for RDF (along with the various view languages that CouchBase and Mongo expose), and it's given me far more insight into RDF than eight years spent trying to understand what the language was all about before then. Indeed, I'd go so far as to say that SPARQL makes RDF, dare I say it, accessible.

For instance, consider one of the more vexing aspects of RDF - how do you deal with composition vs. aggregation? Now, before getting too deep into the realm of modeling, it's worth taking a look at what each of these mean.

In XML, you tend to describe aggregations and compositions the same way - a specific element of one type holds a collection of elements of another type (or subclasses thereof). For instance, a resume element may include a one to many relationship with specific jobs that were held. It may also contain a one to many relationship with the articles, papers or books that you have written. Both of these are called associations - you are associating a given entity with another entity, but they are not quite the same type of thing.

To understand why, you need to ask the question - does the associated object have any meaning outside the context of the containing object? In the case of books, the answer is most assuredly yes - you may not be the only author, the books are likely available by ISBN or on the web, and if you delete the resume, the books do not themselves disappear. In this case, you're dealing with an aggregation - the child entities have  a distinct identity outside the boundaries of the container. In RESTful terms, the child entities are addressable resources.

Composition On the Job

The case of jobs is a little harder. A job is a description of a state - what you were doing at any given time. While the job may have a job title, an associated company and the like, it effectively is a short-hand way of talking about something that you do or did at some point. Take away the context - you - and such jobs generally make much less sense. In a composition, then, the child entities being described are more like states than they are like objects - they are generally only "locally" addressable relative to the containing element or context.

In XML, such structures occur all the time:

<resume>
       <name>Jane Doe</name>
       <job>
             <jobTitle>Architect</jobTitle>
             <company>Collosal Corp.</company>
             <startDate>2011-07</startDate>
             <description>Designed cool stuff.</description>
      </job>
       <job>
             <jobTitle>Senior Programmer</jobTitle>
             <company>Big IT Corp.</company>
             <startDate>2008-05</startDate>
             <endDate>2011- 04</endDate>
             <description>Programmed in many cool languages, and built some stuff.</description>
      </job>
       <job>
             <jobTitle>Junior Programmer</jobTitle>
             <company>Small IT Corp.</company>
             <startDate>2005-05</startDate>
             <endDate>208- 04</endDate>
             <description>Programmed in a couple of cool languages, and built some other stuff.</description>
      </job>
</resume>

So the question here is whether a job is an aggregation or a composition. A good way of thinking about this is to ask yourself whether, if you took turned each job into a separate XML document, whether it has enough context information to make sense:


       <job>
             <jobTitle>Junior Programmer</jobTitle>
             <company>Small IT Corp.</company>
             <startDate>2005-05</startDate>
             <endDate>2008-04</endDate>
             <description>Programmed in a couple of cool languages, and built some other stuff.</description>
      </job>


Here there's no "person" that this belongs to - it could be a job held by anybody (or multiple people at the same time, conceivably). Without that context the information here is insufficient to be useful, save in indicating that someone claims to have worked at a given place. <job> is  clearly a composition relationship.

Now zoom down another level to company:


             <company>Small IT Corp.</company>

Curiously enough, this "document" is actually a stand-alone entity. If I had a database with resumes from a number of different people, being able to see all companies that are represented by the database would be a key requirement, and what's more, as a database designer I'd probably want to insure that there is one and only one such representation of that entity's name to be able to partitioning the data unnecessarily. It's an association.

In RDF (as expressed in Turtle), this becomes much more evident (I'm suppressing namespaces here for ease of reading):
person:Jane_Doe rdf:type class:Person;
          person:name "Jane Doe";
          person:job [
                  job:jobTitle "Architect";
                  job:company company:ColossalCorp;
                  job:startDate "2011-07"^^xs:date;
                  job:description "Designed cool stuff"
                  ],

                  [
                  job:jobTitle "Senior Programmer";
                  job:company company:BigITCorp;
                  job:startDate "2008-05"^^xs:date;
                  job:endDate "2011- 04"^^xs:date;
                  job:description "Programmed in many cool languages, and built some stuff."
                  ],
                  [
                  job:jobTitle "Junior Programmer";
                  job:company company:SmallITCorp;
                  job:startDate "2005-05"^^xs:date;
                  job:endDate "2008- 04"^^xs:date;
                  job:description "Programmed in a couple of cool languages, and built some other stuff."
                  ].

company:ColossalCorp rdf:type class:Company;
             company:companyName "Colossal Corporation".

company:BigITCorp rdf:type class:Company;
             company:companyName "Big IT Corporation".

company:SmallITCorp rdf:type class:Company;
             company:companyName "Small IT Corporation".





If you are new to Turtle, this may take a bit of explaining. An expression like

a  b   c;
   d   e.

is a shorthand for the statements  

a  b  c.
a  d  e.

Similarly, 

a  b  c, d, e.

is a shorthand for the statements  

a  b  c.
a  b  d.
a  b  e.

The first (semicolon) set of statements assume that the subject is repeated but with new predicates and objects, while the second (comma) set of statements assume a common subject and predicate but different objects.

However, what's the [] notation mean? The brackets are indicative of a blank node. A blank node can be thought of as being the equivalent of a composition container in XML  For instance, in XML you have the fragment


       <job>
             <jobTitle>Architect</jobTitle>
             <company>Collosal Corp.</company>
             <startDate>2011-07</startDate>
             <description>Designed cool stuff.</description>
      </job>


The <job> element itself doesn't really serve any purpose beyond indicating that its contents are thematically part of a larger construct. It's not in and of itself a property, but more properly a property "bag". As such, it can be thought of as being (somewhat) analogous to an array - you'd reference it not by name but by position in a language such as Java or JavaScript. (We'll get back to that (somewhat) reference shortly).

Internally, RDF triple stores don't use numbers directly for identifying such entities. Instead, they are treated as blank nodes. A blank node is an internal resource identifier, meaning that, unlike most resources, it can't be directly accessed. Instead, (at least from the standpoint of SPARQL) they are usually referenced indirectly by being determined from the local context. Blank nodes are usually shown in articles as having the syntax _:b0, _:b1, etc. For instance, the resume in the above list could be rewritten as:


person:Jane_Doe rdf:type class:Person;
          person:name "Jane Doe";
          person:job _:b0, _:b1, _:b2.
_:b0      job:jobTitle "Architect";
          job:company company:CollosalCorp;
          job:startDate "2011-07"^^xs:date;
          job:description "Designed cool stuff".
_:b1          job:jobTitle "Senior Programmer";
          job:company company:BigITCorp;
          job:startDate "2008-05"^^xs:date;
          job:endDate "2011- 04"^^xs:date;
          job:description "Programmed in many cool languages, and built some stuff."
_:b2          job:jobTitle "Junior Programmer";
          job:company company:SmallITCorp;

          job:endDate "2008- 04"^^xs:date;
          job:description "Programmed in a couple of cool languages, and built some other stuff.".

Notation-wise, this makes the structure of each job more obvious, but gets a little more complicated referentially. Jane Doe has had three jobs. The jobs each have an internal identifier, but because the jobs make no sense out of the context of being jobs for Jane, So why not use explicit identifiers for each job?


person:Jane_Doe rdf:type class:Person;
             person:name "Jane Doe";
             person:job job:105319, job:125912, job:272421.
job:105319   job:jobTitle "Architect";
             job:company company:CollosalCorp;
             job:startDate "2011-07"^^xs:date;
             job:description "Designed cool stuff".
job:125912   job:jobTitle "Senior Programmer";
             job:company company:BigITCorp;
             job:startDate "2008-05"^^xs:date;
             job:endDate "2011- 04"^^xs:date;
             job:description "Programmed in many cool languages, and built some stuff."
job:272421   job:jobTitle "Junior Programmer";
             job:company company:SmallITCorp;

             job:endDate "2008- 04"^^xs:date;
             job:description "Programmed in a couple of cool languages, and built some other stuff.".
company:ColossalCorp rdf:type class:Company.
             company:companyName "Colossal Corporation".

company:BigITCorp rdf:type class:Company.
             company:companyName "Big IT Corporation".

company:SmallITCorp rdf:type class:Company.
             company:companyName "Small IT Corporation".

From the standpoint of SPARQL, it makes very little difference. For instance, if you wanted to get the name and description of each job a person has, the SPARQL query would look something like this:

select ?personName ?jobTitle ?companyName where {
    ?person   rdf:type      class:Person;
              person:name   ?personName;
              person:job    ?job.
    ?job      job:jobTitle  ?jobTitle;
              job:company   ?company.

    ?company  company:companyName
                            ?companyName.
    }
which would produce the table:
?personName?jobName?companyName
"Jane Doe""Architect""Colossal Corporation"
"Jane Doe""Senior Programmer""Big IT Corporation"
"Jane Doe""Junior Programmer""Small IT Corporation"


The ?job variable, in this case, will carry either the blank node job entry or the defined job in exactly the same manner. So why use a blank node? In most cases, they're used because creating explicit named resource URIs can be a pain, or because there is no real advantage to working with them outside of the resource context from which they're called. Typically, most RDF triple stores actually optimize blank nodes separately, and insure that blank nodes never collide within the data system itself.

Another advantage is that it simplifies both Turtle and RDF-XML notation. Turtle notation can use the [] square bracket approach (and can even nest such brackets within other blank node expressions). RDF-XML on the other hand can represent blank nodes in two modes. In the denormalized (or embedded) form, the @rdf:parseType="resource" expression is used on the container node to indicate that it should be represented from a blank node:


<resume rdf:about="person:Jane_Doe">
       <name>Jane Doe</name>
       <job rdf:parseType="resource">
             <jobTitle>Architect</jobTitle>
             <company>Collosal Corp.</company>
             <startDate>2011-07</startDate>
             <description>Designed cool stuff.</description>
      </job>
      ..

</resume>


The same expression can also be normalized:

<rdf:RDF>
   <resume rdf:about="person:Jane_Doe">
       <name>Jane Doe</name>
       <job rdf:resource="b0"/>
       <job rdf:resource="b1"/>
       <job rdf:resource="b2"/>
   </resume>
   <rdf:Description rdf:nodeID="b0">
       <jobTitle>Architect</jobTitle>
       <company rdf:resource="company:ColossalCorp"/>Colossal Corporation</company>
       <startDate rdf:datatype="xs:date">2011-07</startDate>
       <description>Designed cool stuff.</description>

   </rdf:DescriptiOn>
   ...
</rdf:RDF>


The embedded format actually hints at how one can convert XML into RDF - it mostly involves resource references (i.e., company) and using attributes like rdf:parseType (XML to RDF conversion is food for another post).


Localization with Les NÅ“uds Anonymes


Blank nodes can actually be very useful for dealing with groups of properties that differ based upon language or locale. RDF utilizes the @lang expression at the end of strings to handle localization, which can work in simple cases, but if there are a number of properties that are all localized (such as prices and currencies) then blank nodes may be a better solution. For instance,


book:The_Art_Of_SPARQL book:bookLocal
     [
        bookLocal:lang "EN";
        bookLocal:locale "US";
        bookLocal:title "The Art of Sparql";
        bookLocal:price "29.95";
        bookLocal:currency "USD";
     ],

     [
        bookLocal:lang "EN";
        bookLocal:locale "UK";
        bookLocal:title "The Art of Sparql";
        bookLocal:price "21.95";
        bookLocal:currency "GBP";
     ],
     [
        bookLocal:lang "DE";
        bookLocal:locale "DE";
        bookLocal:title "Die Kunst des SPARQL";
        bookLocal:price "24.95";
        bookLocal:currency "EUR";
     ].


In this case, the implied bookLocal objects are blank nodes. You can then retrieve the title, price and currency of a book in a different locale if the US book title is known:

select ?title ?lang ?price ?currency where {
    bind ("DE" as ?locale)
    bind ("The Art of SPARQL" as ?usTitle)
    ?book book:bookLocal [
              bookLocal:title ?usTitle;
              bookLocal:locale "US"].
    ?book book:bookLocal [
              bookLocal:locale ?locale;
              bookLocal:title ?title;
              bookLocal:lang ?lang;
              bookLocal:price ?price";
              bookLocal:currency ?currency".
              ].
    }

This produces the results:
 
?title?lang?price?currency
"Die Kunst des SPARQL""DE""24.95""EUR"

This information can, depending upon the SPARQL engine, be output as JSON or XML. This approach can also be useful for dealing with triples where the object is itself an XML node (such as XHTML content) where different languages are involved, since the XMLLiteral type can't directly take a @lang type extension. I use such blank node assignments quite frequently for precisely that kind of issue, making data modeling when dealing with localization considerably easier.

It's also worth noting that even though the nodes themselves are blank (or anonymous), that doesn't mean that they can't be associated with schematic types. For instance, in the previous example, the inclusion of a single statement in each block means that you can use RDFS/OWL validation on each "bookLocal" object:

 book:The_Art_Of_SPARQL book:bookLocal
     [
        rfd:type class:bookLocal;
        bookLocal:lang "EN";
        bookLocal:locale "US";
        bookLocal:title "The Art of Sparql";
        bookLocal:price "29.95";
        bookLocal:currency "USD";
     ],

These basically correspond to anonymous objects within JavaScript, with the added benefit that with RDF you can still do all of the type constraint checks that you could do with XSD (and more).


The Edge of the Graph


One final note about blank nodes - they are very useful for establishing the boundaries between objects for purposes such as getting the result of a Describe or deleting distinct objects (rather than just individual assertions) from a database. The describe statement, when given a subject, follows each assertion from the initial subject to each object. If the object is an atomic value, the assertion is kept, but no further search is done. If the object is an object URI, the same thing happens.

However, if the object is a blank node, then the blank node is added as a subject and the same process is used. Deleting an object involves doing a "Describe" on the resource, finding all assertions for which the subject is some other assertion's object, then deleting all of these links.

This can be especially useful for handling updates into RDF Databases. I'll be covering this in more details in a later post.


Nothing Much? Nah!


Blank nodes are a useful feature of RDF, are well supported by SPARQL and Turtle notation, and can help to differentiate between aggregate structures (references to external objects within the system) and composed structures (references to internal entities that only make sense within the context of a given external entity). They are used heavily by both Turtle and RDF-XML parsers, and they can both be used to define the boundaries of resources within the system and delete those resources in a logical and consistent fashion. They should be seen as indispensable tools of the working ontologist. 

Semantics + Search : MarkLogic 7 Gets RDF


Let's talk Org Charts for a moment. Everyone knows what an org chart looks like. At the top you have the boss guy, Mr. Bigg, CEO of Bigg Business, Inc.

The Head Honcho

At the second tier, you have his lieutenants, the guys that head up the various functions within the organization (flipping to show relationships more clearly).





What emerges as part of this is the fact that you have what appears to be a tree. This can be made even more obvious when you start jumping to the next level (just showing a couple of people from the programming department):



If asked at this point, what would be the best way to store information about this organization chart, chances are pretty good that you'd opt for a language like XML or JSON, since there is a clear container/contained relationship between all the parties. However, suppose for a moment that the IT programmers are intended to be support personnel. While Ian Geek signs their paycheck, they actually work in different departments. Then you end up with something like this.


sem3.dot


Oops.

What had been a nice, clean hierarchical chart suddenly gets considerably more complicated. There are actually three things worth noting here - first, that we've jumped from having one relationship - "reports to" - to having two - "reports to" and "assists". The second is that the "assists" relationship does not in fact support a hierarchical distribution at all - Jane supports both Owen Munny and Bartholomew Bigg, who are ostensibly at different levels within the organization. A side note (and the number three thing in the list) is that whereas the "reports to" relationship is always one to one, that's not true of "assists" - Jane assists two people.

This is the reason why data designers and information architects get the big bucks - the real world is usually considerably more messy than a hierarchy. If you had a SQL database, modeling this relationship wouldn't be all that hard - every person gets a primary key that identifies them, the reports to relationship would incorporate a foreign key directly to that person's manager, and the assists relation involves a pointer to a table which then maps the primary key to zero or more associated keys.

However, there may in fact be any number of reasons why you would prefer to keep your data in XML (or in JSON, the same arguments apply there). Hierarchies are convenient - they provide a means for organization information in a consistent fashion, they provide levels of abstraction for collections, and they are far more useful for transferring information than linear tables. As long as you're dealing with properties, or bags of properties, they work very nicely indeed.

The problem comes when you start dealing with discrete objects. In an organization, a person is a discrete object. A person may have one or more names, one or more job titles, one or more locations, one or more responsibilities. Some of these are simple properties - a name or a job title is a simple text string, for instance, the start (and potentially end) date are, well, dates. Other properties get a little more complex, and really depend upon what specifically is being modeled.

Locations are good examples of this - most organizations have rooms (or at least cubicles) where specific people work, and a way of identifying that  room via some form of key. The location is in fact a "thing" - a resource - it is in a certain building, specific floor, and may have its own phone number.  It makes sense for the person to have a relationship with that location, but that relationship isn't necessarily a containment or ownership relationship. An room may have more than one person in it, or be a temporary cubicle for contractors. A person may have multiple places they work from, including possibly from home.

Again, all of this can be modeled. However, there's something of a twist here. Suppose that in your organization, one database holds the names, positions, salaries and other business related information about a person. Another database holds the locations and names of people who are at those locations.

And then a reorganization occurs. You are tasked with identifying the org chart relationships then, when possible, moving people so that they are closest first to the people that they assist, then, when possible to the people that manage them. In an organization with 50 people, this can be a time consuming activity. In an organization with 10,000 people, it becomes a logistical nightmare. I should note here that a great number of business problems (and opportunities) share the same basic characteristics. I'm laying this one out just as an example.

There are actually several key issues that are brought up here. The first is the fact that most organizations have real problems with identity management - each database has a different (usually numeric) key that it internally uses to identify a person, location, department and so forth, and as a consequence finding whether two records refer to the same person typically comes down to having to identify sets of common characteristics, and increases the likelihood of error (as well as introduces computational costs).

Getting a new database to match keys by itself isn't enough (it in fact simply defers the problem for a bit) - you essentially need to identify and inventory the resources that your organization has, and then assign to each of them a unique ID - a GUID, or better, a Uniform Resource Identifier (URI) (also referred to as a uniform resource number (URN) or an international resource identifier (IRI), though each have subtly different meanings.  A URI is a unique string that globally identifies a resource, whether that resource is a book, a web page, a person, or a room/cubicle. It may look something like schema://biggbusiness.com/person/jane_doe, or may be more cryptic (such as urn:biggbiz:1295102:39302185:1593). What's important here is that it is effectively unique - it not only identifies the person within a single database, but it identifies that person (or other resource) globally.

Once you have that globally unique identifier, then and only then can you start attaching other identifiers that may not be global. In many respects, this is the core of "semantics":  Uniquely identify the common resources (and resource types) in an organization, use those resource keys in what amounts to a "columnar" database or graph store, establish the relationships that exist between these resources then associate enough properties and local identifiers to these global identifiers to make search feasible.
SPARQL is a big part of that semantic layer. SPARQL is to Semantic "Triple Stores" what XQuery is to XML, JavaScript (and associated JSON query dialects) is to JSON stores and SQL is to relational data. It allows you to query the relationships between these global objects, returning in turn either a binary answer, a table of variable values or other triples, depending upon what's asked, and most SPARQL engines also have a mechanism to create (simple) JSON and XML output structures.

Triple stores (named because columnar stores can be broken down into a three value set consisting of a "subject", a "predicate" property, and an "object" URI or scalar value) and SPARQL together are consequently useful because they allow you to perform joins across relationships on objects, even when the objects being joined are not simple tables. The keys they are joining on are URIs, not sequential indexes, and they transcend any single database.

This really comes in handy when dealing with controlled vocabulary lists. Suppose that I wanted to get a list of all people that directly report to Ian M. Geek and I know his URI. With SPARQL I can retrieve all people who have a reports to association with Ian (of course, I can do that with XQuery as well), since it's simply a string match. However, what I can also do is retrieve all people who report directly or indirectly with Ian - they report to someone who reports to someone who reports to him, for instance. This is what's called chaining in semantics and is one of the things that makes SPARQL very useful. Moreover, if I wanted to get the names and titles of everyone who is in Ian's reporting chain, with XQuery I would have to retrieve the document for each person and then get the name and title property, while with SPARQL, I'm dealing simply with assertions - an object may be described as the set of all triples that have the same identifier URI rather than an explicit "document", and so the SPARQL processor need only look at specific relationships, not all of them.

Having said that, SPARQL and Triple Stores in general are not ideal for other kinds of queries. Suppose, for instance, that I want to find out how reports to the CTO, but I don't know the CTO's name. There's actually two problems here - first, find a string match in an (unspecified) field, possibly with a regular expression, which will then retrieve the identifier of the corresponding person(s), then perform the previous query. 

Triple stores do provide lexical search capabilities and indexing, but in general these are expensive operations, far more than is the case for XML stores.  However, this is a classic "search" problem, and finding whole or near matches of terms is something that can be handled quite easily.

Moreover, SPARQL doesn't handle much beyond the semantic search - you can transform it with XSLT or similar applications, but these kinds of activities are often somewhat limited in scope. XQuery, on the other hand, is nearly as robust a mechanism for creating output content as it is for search.

Because of this, building a SPARQL application almost invariably has required invoking a triple store via a web service from some controlling language then pulling the result back and transforming it to meet the particular needs. Since you're sending data "over the wire" this has performance implications, and keeping the database in synchronization requires some serious headstands. However, that's now changed.

The MarkLogic 7.0 server was announced at MarkLogic World 2013 last month, becoming the first XQuery database to incorporate a semantics index and a SPARQL layer.  A bit of disclosure here - Avalon Consulting, LLC,  works as a primary partner for MarkLogic, largely because we value their product quite highly in the development of enterprise information solutions. I've also been lobbying MarkLogic to develop a SPARQL layer for more than three years, so I was absolutely giddy to discover that they had finally gone ahead and done it.

Because this technology is fairly complex, they are implementing it in stages, with Marklogic 7.0 supporting the SPARQL 1.0 standard and several low hanging fruit in SPARQL 1.1. They will then implement the balance of the SPARQL 1.1 layer (including SPARQL UPDATE) in a subsequent release, and may include some of the more sophisticated support for the OWL 2.0 web ontology standard (commonly referred to as the inferencing layer).

I recently (finally) had a chance to review the MarkLogic Server 7.0 Early Access 2 release, after having some trepidation that it would be too little too late. After at least a preliminary analysis, I think I can safely say that if and when MarkLogic ever goes public, I would buy as much of their stock as possible. Even without the full 1.1 support, it supports the SPARQL 1.0 standard remarkably well, it is as fast as I've come to expect MarkLogic products to be (that is, very), and most of what I personally would like to do with RDF I can, albeit not necessarily directly through SPARQL.

I can also do the kind of queries I was talking about above, and combine search and semantic queries into a single, internal operation that is breathtakingly fast (in great part because it is NOT having to go out to an external server). For instance, in the above queries I would use XQuery to do the text search to get those documents that include the search term in a specific set of fields (looking for CTO, for instance), then will pass these documents, with their associated triples to a SPARQL query. This reduces the search dramatically from comparing against the overall index to comparing against perhaps a few dozen entries. In SPARQL, such reductions are what makes working against tens of billions to trillions of triples feasible. These results, in turn, are filtered to retrieve name and job title from those working for the CTO, and then are returned as either JSON structures, custom XML, or generic map structures, depending upon need. These can then be transformed with XQuery into everything from HTML to SVG to PDFs or spreadsheets.



The real value here then comes in the ability to combine, transform and cache such output, as well as to manage processing. If a person quits, for instance a workflow would be initiated upon state change (through a process called a reverse query) that would query all locations that are associated with a given person (has an "claims" pointer, for instance) and would free those up. A similar query can determine whether there are any people who have placed reservations upon that room should it become freed, and will associate the room to the person based upon test criteria (the person who has the oldest reservation, for instance). This becomes the foundation for a rules based orchestration system that can replace a complex command and control system. Because such systems often have a significant number of inter-dependencies on different resources, a semantic system is actually preferable in this regard.

To sum it up, MarkLogic with semantics has what it takes to learn, to create associations where none previously existed, to judge likelihoods (I'm salivating at the prospect of writing a Bayesian analysis parser on top of it), to generate user interfaces that change in response to what it needs, and to commit to conclusions and actions based upon its own internal analysis. I'm noting that a number of entity enrichment, business intelligence and natural language processing companies are migrating to MarkLogic as a foundational platform for their own offerings, and I fully anticipate this trend to increase dramatically as the combination of XQuery, SPARQL and SQL (yup, it does support that too, thank you Mary Holstege!) makes MarkLogic a nexus  for the intelligent machine.

Monday, April 22, 2013

Age Discrimination or Reasonable Expectations?


I recently saw a discussion about Age Discrimination in the IT Workforce (http://www.linkedin.com/today/post/article/20130422020049-8451-the-tech-industry-s-darkest-secret-it-s-all-about-age) and it got me thinking about it beyond the immediate knee-jerk reaction of fear about my own job standing. In general, while IT people tend to have an outsized impression of the value of their code within an organization (sometimes deserved, sometimes not) from an organizational perspective most code is simple depreciating assets that will need to be replaced over time. For that reason, the cost issue of paying a senior developer at 55 earning 2-3 times a junior programmer's salary for what is perceived as decaying assets begins to make sense.

Moreover, there are typically several career tracks which employers look for with technical people, and these are very much tied into age and experience. Looked at in that perspective, the age distribution within the industry begins to look a little bit more rational, but it also means that you should be thinking about career management from the time you leave college.

At 25, you're basically an apprentice - learning the "art" of programming, gaining experience, putting in the hours to prove yourself. You don't have a family, you're willing to work 80 hours for a 40 hour a week job, coding is still new and shiny, and the code that you write, while probably not brilliant, will likely be more innovative simply because you're not locked into specific patterns yet. Employers will hire you because they are generally less concerned with quality than they are costs. You're expendable, and for this age group being expendable is not necessarily a bad thing - it gives you exposure to different problem domains, and teaches you how to stay nimble in the market.

By 40, you're probably a good programmer, skilled in several languages, with a few major project successes and failures under your belt. However, your value to the organization increases if you can use that experience to train up, mentor and manage the younger programmers, learn how to knock heads to keep egos out of the way of making deadlines, or to transfer that experience more into architecture or product design. It's also a good time to go journeyman - gain experiences as a consultant, learn how to work directly with clients and how to recognize when problems are better handled with social programming than technological ones. Start working at the systems architecture level. Write a book or three.

As you head into your fifties, it's a better time to start a company, to expand your consulting to large companies or to go into research - it's also not a bad time to go back to school and get that PhD that you didn't have time for earlier, since kids are nearing adulthood by then if not already there, you have a much better breadth of knowledge on which to base your studies and you are far more likely to get into a faculty position or a research institute if you're seasoned with real world experience (and you'll be a better teacher to your students as well).

From an entrepreneurial standpoint you have the connections that are so critical to starting a business and building up a market. As a consultant, you don't have the demands of home and hearth to hold you back any more to the same extent you did when your kids are still young, so you are capable of travel and long distance gigs. You're also more likely to be acting in a mixed business/technical perspective, providing architectural or business guidance while some whiz kid in his twenties bangs out the cool code.

If you're 55 and still doing nothing but coding, senior management will wonder why they are shelling out so much for that code, no matter how brilliant, since they see the code as being ultimately expendable.

I'm fifty. I am a consulting information architect, have written more than a dozen books, and regularly consult with federal agencies and Fortune 100 companies. I still write code, but most of that code anymore gets put into my books or articles or becomes proof of concepts for clients when trying to sell a new technology idea. In other words, the coding is simply a tool in my suite of tools now, The pay is generally better, the ability to influence projects is considerably more significant, and I find it easier to deal directly with the C-level executives that make the decisions.

What this means to a young programmer in particular is simple - expand your horizons. The code is important, but it is not the only important piece in the puzzle, and in many respects no matter how good a programmer you are, it's your ability to navigate the social shoals and reefs that determine how successful you are in your career. As an employer, I'd be as skeptical about a 55 year old programmer as I would be about a 25 year old information architect - neither one has the experience that I need for the jobs at hand.

Sunday, April 7, 2013

Book Chapter and Verse - HTML5 + RDFa

Some time back I did some work for a religious publisher, and at the time began to play around with the question of how we could take a lot of the content that they had for their book publications and make them more accessible, especially given the interrelatedness of their various content. I'd begun playing at the time with RDF, and the notion of building graphs for solving this particular problem was attractive, but the tools were simply not yet there to do more than contemplate.

Recently, I had a chance to talk with a friend I'd first met there, and I got to thinking about the idea of making use of HTML + one of the less discussed aspects of the HTML5 revolution - the utilization of RDF for Attributes, otherwise known as RDFa. The idea behind RDFa is relatively simple - put into attributes sufficient information to be able to embed semantic markup into an HTML5 document, then once you have this, converting this RDFa into RDF via the use of a transformation technology called GRDDL (which is typically implemented in XSLT, but can also be done in Java).

Since I've been doing some exploration of semantics for a couple of large media clients, I decided it would be a good time to experiment, and get a better handle on working with RDFa -- including putting together some more challenging constructs in order to see how this can be effectively integrated into RDF/OWL.

What I ultimately came up with was a very tiny book of "holy writ" that may actually end up becoming a short story at some point (never throw away good concepts). At the moment, this exists primarily to show concepts, but I suspect I'll be able to build more up. The original XHTML is shown in Listing 1.

<html
    xmlns="http://www.w3.org/1999/xhtml">
    <body>
        <header>
            <a name="Michael"><h1>The Book of Michael</h1></a>
        </header>
        <article>
            <a name="Michael-1-1" id="Michael-1-1" >                <p>In the heart of the world, not long after the beginning, a man named Michael Grenadine, he who was known as The Librarian, walked out of the Wastes.</p></a>
            <a name="Michael-1-2" id="Michael-1-2" >
                <p>Travelled did he with naught but a humble mule, which he called The Horse With No Name, verily to wreak confusion and wonder among the unenlightened.</p></a>
            <a name="Michael-1-3" id="Michael-1-3">                <p>Within his wagon, hauled by his humble mule, did Michael carry with him many books, written by the Monks of Knowledge and the Sisters of Understanding, the better to preserve the knowledge and wisdom of ages past, as well as the cautionary tales that had led to the Wastes.</p></a>
            <a name="Michael-1-4" id="Michael-1-4">
<p>Weary was he, and sore afoot, for he had travelled many days and nights, had wandered through mountains and deserts, searching for a new home, until finally he collapsed, ill and feverish.</p></a>
  <a name="Michael-1-5" id="Michael-1-5">
                <p>There would Michael have died, had his mule not begun to bray in fear for his master, for its master was a kind man who fed the mule even when he himself had nothing, and groomed his mule before ever taking himself down to sleep.</p></a>
            <a name="Michael-2-1" id="Michael-2-1">
                <p>Not far away, near the town of Seattlec'l, a young woman named Alara Dishaean was tending her cows when she heard the braying of a mule.</p></a>
            <a name="Michael-2-2" id="Michael-2-2">
                <p>Curious, Alara left her farm on a Bantu Horse until she entered into the Great Forest, where she found Michael upon the forest floor, unconscious.</p></a>
        </article> 
    </body>
</html>
Listing 1. The original HTML for the Book of Michael.

Overall, there are a number of distinct types of entities at work here - concepts, people, locations, along with the implied book, chapter and verse that one might expect from a work of this sort. It takes a bit of data modeling to figure out what specifically you want to track (we'll come back to that later), but ultimately what tends to emerge is that you want to maintain one "namespace" for each type of object that you'll want to model. RDFa 1.0 introduced the concept of the curie, but because the use of namespace prefixes are somewhat problematic, RDFa 1.1 for HTML introduced the prefix attribute, which is typically placed on the outermost viable container, and consists of the form:

prefix="prfx1:  http://www.example.org/prefix1 prfx:2 http://www.example.org/prefix2 ..." 

and so forth. These prefixes reduce the overall size of the attributes within the HTML document, considerably, and make them marginally more legible. In the case of the example document, the new html header now looks as follows:

<html xmlns="http://www.w3.org/1999/xhtml" vocab="http://OrderOfTheBook.org/xmlns/verse#"
    prefix="bio: http://OrderOfTheBook.org/xmlns/bio# 
    location: http://OrderOfTheBook.org/xmlns/location# 
    concept: http://OrderOfTheBook.org/xmlns/concept#
    writ: http://OrderOfTheBook.org/xmlns/writ#
    verse: http://OrderOfTheBook.org/xmlns/verse#
    chapter: http://OrderOfTheBook.org/xmlns/chapter#
    book: http://OrderOfTheBook.org/xmlns/book#
    class: http://OrderOfTheBook.org/xmlns/class#
    dc: http://purl.org/dc/terms/"
    about="chapter:Michael" typeof="class:Chapter">



The @prefix attriibute identifies all of the CURIE namespaces and their associated prefixes. Each is essentially a different vocabulary of terms.

The @about attribute identifies that this document is a description of the chapter "Michael". The expression chapter:Michael is a prefix qualified name, and is equivalent to http://OrderOfTheBook.org/xmlns/chapter#Michael
and is effectively the "subject" of this particular document. Unless it is overridden in an internal content, this is the context that all properties belong to.

The advantage of working with CURIES should be evident there - they make the code considerably easier to read.

The @typeof attribute is a reference to an RDF class, in this case "class:Chapter". This is actually a pretty significant innovation, because what we have done is defined semantically that this particular construct is a Chapter, which means that rules and logic that apply to chapters apply here as well. Moreover, there is now a different semantic markup beginning to emerge on top of the HTML semantics, which don't really even have a notion of chapter.

The @vocab attribute identifies what the default semantic namespace is, and as
so is inherited by subordinate containers until a container has a different @vocab defined, at which point the innermost vocab becomes the default. In this case,

After the <html> element, the next element is the header, which contains the title of the chapter, as shown in Listing 2.

    <body>
        <header>
           <h1 property="dc:title">The Book of Michael</h1>
        </header>
Listing 2. An RDFa property.


This uses the Dublin Core namespace to identify the title of the document, which internally would end up creating a triple that looks like:


@prefix chapter: <http://OrderOfTheBook.org/xmlns/chapter#> .
@prefix dc: <http://purl.org/dc/terms/> .
chapter:Michael     dc:title   "The Book of Michael".

in Turtle notation.

The next section identifies for the book the first chapter (listing 3):


        <span inlist="" property="book:chapter" resource="chapter:Michael-1"/>
        <article about="chapter:Michael-1" typeof="class:Chapter">
            <header>
                <h2 property="dc:title">Michael: Chapter 1</h2>
            </header>
Listing 3. Defining the chapter.

This is a bit more complicated - the first <span> statement is within the context of the global book, <book:Michael>, and asserts that there is a property called book:chapter with a link to the chapter resource URI, <chapter:Michael-1>. The inlist attribute (with an empty value in XHTML or just the attribute without a value at all for HTML) requires some explanation. Ordinarily, there is no concept of sequential ordering in RDF, but ordered sequences do occur in real life. To get around this, RDF defines the notion of a list. that uses blank nodes and several key RDF properties. The @inlist attribute signals to the parser that the resource should be attached in a list sequence to the context object using the specified property predicate (in this case the property <book:chapter>). This means that the book will have an Turtle notation that looks like:

book:Michael book:chapter (chapter:Michael-1 chapter:Michael-2 chapter:Michael-3). 

and so forth.

The next line identifies the article to be "about" <chapter:Michael-1>, making this the new context that everything else is related to as subject. This also identifies this article as being of the Chapter type. This can be useful to create specific visual identities for the various object types in your system because you can create a CSS rule such as

article [typeof=class:Chapter] {background-color:lightBlue;font-family:Arial; ...}

that would provide a visual rendering of any Chapter object, regardless of the underlying HTML semantics.

The <header><h1> elements also provide a label for the new chapter object (remember, we're now in the chapter context, not the book). Note that when you have a @property attribute with neither content nor a resource, then the string representation of the contents of that attribute's element becomes the value used for the property association (here, the Dublin Core <dc:title> property). This holds true even if the element has child elements within it, a fact we take advantage of at the verse level.

Each verse follows a similar convention as shown in Listing 4.


         <div inlist="" property="chapter:verse" resource="verse:Michael-1-1">
            <a name="Michael-1-1" id="Michael-1-1" about="verse:Michael-1-1" typeof="class:Verse">
                <p property="text"><span property="dc:title" content="Michael 1-1"/>In the heart of the world, not long after the beginning, a man named <span property="bio" resource="bio:Michael_Grenadine">Michael Grenadine</span>, he who was known as <span property="concept" resource="concept:librarian">The Librarian</span>, walked out of <span property="location" resource="location:The_Wastes">the Wastes</span>.</p>
            </a>
          </div>
Listing 4. The HTML verse content. 


Here, you have a <div> element that identifies the resource in question as being part of the <chapter:verse> list. Within this you have the identifier itself (bound with an <a name=".."> element) along with its class indicator. Everything within this element is now part of the verse scope.

The paragraph <p> element also includes a property, but in this case it is given only as a local name. The property="text" statement takes advantage of the default namespace that was defined in the header, which defines the default semantic namespace as being the verse: namespace. What this means in practice is that the expression
<p property="text">
is actually a shorthand notation for
<p property="verse:text">,
which, since the element has neither a @content or a @resource attribute, means that the text string of the content, minus any internal markup, is the object of the assertion:
verse:Michael-1-1   verse:text    """In the heart of the world, not long after the beginning, a man named Michael Grenadine, he who was known as The Librarian, walked out of the Wastes.""".

The triple quotes are multi-line quotes, used both to hold content that may span multiple lines and used to safely encapsulate both single and double quotes which can cause problems with text parsing logic.

Within the verse: namespace, there are a number of organizational concepts - verse:loc for locations, verse:bio for biographical entities, verse:concept for conceptual entities (love, war, illness, death), terms that you may expect with a semantic knowledge ontology system. These keywords can often provide multiple ways of simultaneously organizing content, either in hierarchical taxonomies or in more freeform associational structures, but they also can make finding related topics far easier.

The full structure for the HTML document is given in Listing 5.


<html xmlns="http://www.w3.org/1999/xhtml" vocab="http://OrderOfTheBook.org/xmlns/verse#"
    prefix="bio: http://OrderOfTheBook.org/xmlns/bio# 
    location: http://OrderOfTheBook.org/xmlns/location# 
    concept: http://OrderOfTheBook.org/xmlns/concept#
    writ: http://OrderOfTheBook.org/xmlns/writ#
    verse: http://OrderOfTheBook.org/xmlns/verse#
    chapter: http://OrderOfTheBook.org/xmlns/chapter#
    book: http://OrderOfTheBook.org/xmlns/book#
    class: http://OrderOfTheBook.org/xmlns/class#
    owl: http://www.w3.org/2002/07/owl#
    rdf: http://www.w3.org/1999/02/22-rdf-syntax-ns#
    rdfs: http://www.w3.org/2000/01/rdf-schema#
    dc: http://purl.org/dc/terms/"
    about="book:Michael" typeof="class:Book">
    <body>
        <header>
           <h1 property="dc:title">The Book of Michael</h1>
        </header>
        <span inlist="" property="book:chapter" resource="chapter:Michael-1"/>
        <article about="chapter:Michael-1" typeof="class:Chapter">
            <header>
                <h2 property="dc:title">Michael: Chapter 1</h2>
            </header>
            <span inlist="" property="chapter:verse" resource="verse:Michael-1-1"/>
            <a name="Michael-1-1" id="Michael-1-1" about="verse:Michael-1-1" typeof="class:Verse">
                <p property="text"><span property="dc:title" content="Michael 1-1"/>In the heart of the world, not long after the beginning, a man named <span property="bio" resource="bio:Michael_Grenadine">Michael Grenadine</span>, he who was known as <span property="concept" resource="concept:librarian">The Librarian</span>, walked out of <span property="location" resource="location:The_Wastes">the Wastes</span>.</p>
            </a>
            <span inlist="" property="chapter:verse" resource="verse:Michael-1-2"/>
            <a name="Michael-1-2" id="Michael-1-2" about="verse:Michael-1-2" typeof="Verse">
                <span property="dc:title" content="Michael 1-2"/>
                <p property="text">Travel did he with nought but a humble <span property="concept" resource="concept:The_Mule_Of_Michael">mule</span>, which he called <span property="concept" resource="concept:The_Horse_With_No_Name">The Horse With No Name</span>, verily to wreak confusion and wonder among the <span property="concept" resource="concept:Unenlighted">unenlightened</span>.</p>
            </a>
            <span inlist="" property="chapter:verse" resource="verse:Michael-1-3"/>
            <a name="Michael-1-3" id="Michael-1-3" about="verse:Michael-1-3" typeof="Verse">  
                <p property="text"><span property="dc:title" content="Michael 1-3"/>Within his wagon, hauled by his humble <span property="concept" resource="concept:The_Mule_Of_Michael">mule</span>, did Michael carry with him <span property="concept" resource="concept:The_Library_Of_Michael">many books</span>, written by the <span property="concept" resource="concept:Monks_Of_Knowledge">Monks of Knowledge</span> and the <span property="concept" resource="concept:Sisters_Of_Understanding">Sisters of Understanding</span>, the better to preserve the knowledge and wisdom of ages past, as well as the cautionary tales that had led to the creation of <span property="location" resource="location:The_Wastes">the Wastes</span>.</p>
            </a>
            <span inlist="" property="chapter:verse" resource="verse:Michael-1-4"/>
            <a name="Michael-1-4" id="Michael-1-4" about="verse:Michael-1-4" typeof="Verse">                
                <p property="text"><span property="dc:title" content="Michael 1-4"/><span property="concept" resource="concept:Weariness">Weary was he, and sore afoot,</span> for he had <span property="concept" resource="concept:Travel">travelled many days and nights</span>, had wandered through <span property="concept" resource="concept:Mountain">mountains</span> and <span property="concept" resource="concept:Desert">deserts</span>, searching for a new home, until finally <span property="concept" resource="concept:Illness">he collapsed, ill and feverish</span>.</p>
            </a>
            <span inlist="" property="chapter:verse" resource="verse:Michael-1-5"/>
            <a name="Michael-1-5" id="Michael-1-5" about="verse:Michael-1-5" typeof="Verse">                
                <p property="text">There would Michael have died, had his <span property="concept" resource="concept:The_Mule_Of_Michael">mule</span> not begun <span property="concept" resource="concept:Loyalty">to bray in fear for his master</span>, for its master was a <span property="concept" resource="concept:Kindness">kind man who fed the mule even when he himself had nothing</span>, and <span property="concept" resource="concept:Caring">groomed his mule before ever taking himself down to sleep</span>.</p>
            </a>
        </article>
        <span inlist="" property="book:chapter" resource="chapter:Michael-2"/>
        <article about="chapter:Michael-2" typeof="class:Chapter">
            <header>
                <h2 property="dc:title">Michael: Chapter 2</h2>
            </header>
            <span inlist="" property="chapter:verse" resource="verse:Michael-2-1"/>
            <a name="Michael-2-1" id="Michael-2-1" about="verse:Michael-2-1" typeof="class:Verse">
                <p property="text"><span property="dc:title" content="Michael 2-1"/>Not far away, near the town of <span property="location" resource="location:Seattle_Delshaean_Era">Seattlec'l</span>, a young woman named <span property="bio" resource="Alara_Dishaean">Alara Dishaean</span> was tending her <span property="concept" resource="concept:Cattle">cows</span> when she heard <span property="concept" resource="concept:Mule">the braying of a mule</span>.</p>
            </a>
            <span inlist="" property="chapter:verse" resource="verse:Michael-2-2"/>
            <a name="Michael-2-2" id="Michael-2-2" about="verse:Michael-2-2" typeof="class:Verse">
                <p property="text"><span property="dc:title" content="Michael 2-2"/><span property="concept" resource="concept:Curiosity">Curious,</span> Alara left her farm on a <span property="concept" resource="concept:Bantu_Horse">Bantu Horse</span> until she entered into the <span property="location" resource="location:Great_Forest">Great Forest</span>, where she found <span property="bio:" resource="bio:Michael_Grenadine">Michael</span> upon the forest floor, unconscious.</p>
            </a>
        </article>
    </body>
</html>
Listing 5. The full HTML-based book.

Up to now, this article has focused on the construction of the RDFa, but has not yet answered how this gets translated into RDF. The answer to this is to make use of a program called an RDFa Parser/Distiller. The one I've used for these examples is available online as a Python application at http://www.w3.org/2012/pyRdfa/ . This runs as a service, and lets you pass RDFa as a text stream, upload a file, or parse an online page with embedded RDFa code. Overall pyRdfa seems to offer the most comprehensive RDFa 1.1 coverage.

Running the above example through the parser, specifying for XHTML5+RDFA 1.1 input and Turtle output (see Figure 1), pyRDFa produces the Turtle triples given in Listing 6.

Figure 1. The configuration screen for pyRDFa text input.

@prefix bio: <http://OrderOfTheBook.org/xmlns/bio#> .
@prefix book: <http://OrderOfTheBook.org/xmlns/book#> .
@prefix chapter: <http://OrderOfTheBook.org/xmlns/chapter#> .
@prefix class: <http://OrderOfTheBook.org/xmlns/class#> .
@prefix concept: <http://OrderOfTheBook.org/xmlns/concept#> .
@prefix dc: <http://purl.org/dc/terms/> .
@prefix location: <http://OrderOfTheBook.org/xmlns/location#> .
@prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> .
@prefix rdfa: <http://www.w3.org/ns/rdfa#> .
@prefix verse: <http://OrderOfTheBook.org/xmlns/verse#> .

<> rdfa:usesVocabulary verse: .

book:Michael a class:Book;
    book:chapter ( chapter:Michael-1 chapter:Michael-2 );
    dc:title "The Book of Michael" .

chapter:Michael-1 a class:Chapter;
    chapter:verse ( verse:Michael-1-1 verse:Michael-1-2 verse:Michael-1-3 verse:Michael-1-4 verse:Michael-1-5 );
    dc:title "Michael: Chapter 1" .

chapter:Michael-2 a class:Chapter;
    chapter:verse ( verse:Michael-2-1 verse:Michael-2-2 );
    dc:title "Michael: Chapter 2" .

verse:Michael-1-1 a class:Verse;
    verse:bio bio:Michael_Grenadine;
    verse:concept concept:librarian;
    verse:location location:The_Wastes;
    verse:text "In the heart of the world, not long after the beginning, a man named Michael Grenadine, he who was known as The Librarian, walked out of the Wastes.";
    dc:title "Michael 1-1" .

verse:Michael-1-2 a verse:Verse;
    verse:concept concept:The_Horse_With_No_Name,
        concept:The_Mule_Of_Michael,
        concept:Unenlighted;
    verse:text "Travel did he with nought but a humble mule, which he called The Horse With No Name, verily to wreak confusion and wonder among the unenlightened.";
    dc:title "Michael 1-2" .

verse:Michael-1-3 a verse:Verse;
    verse:concept concept:Monks_Of_Knowledge,
        concept:Sisters_Of_Understanding,
        concept:The_Library_Of_Michael,
        concept:The_Mule_Of_Michael;
    verse:location location:The_Wastes;
    verse:text "Within his wagon, hauled by his humble mule, did Michael carry with him many books, written by the Monks of Knowledge and the Sisters of Understanding, the better to preserve the knowledge and wisdom of ages past, as well as the cautionary tales that had led to the creation of the Wastes.";
    dc:title "Michael 1-3" .

verse:Michael-1-4 a verse:Verse;
    verse:concept concept:Desert,
        concept:Illness,
        concept:Mountain,
        concept:Travel,
        concept:Weariness;
    verse:text "Weary was he, and sore afoot, for he had travelled many days and nights, had wandered through mountains and deserts, searching for a new home, until finally he collapsed, ill and feverish.";
    dc:title "Michael 1-4" .

verse:Michael-1-5 a verse:Verse;
    verse:concept concept:Caring,
        concept:Kindness,
        concept:Loyalty,
        concept:The_Mule_Of_Michael;
    verse:text "There would Michael have died, had his mule not begun to bray in fear for his master, for its master was a kind man who fed the mule even when he himself had nothing, and groomed his mule before ever taking himself down to sleep." .

verse:Michael-2-1 a class:Verse;
    verse:bio <Alara_Dishaean>;
    verse:concept concept:Cattle,
        concept:Mule;
    verse:location location:Seattle_Delshaean_Era;
    verse:text "Not far away, near the town of Seattlec'l, a young woman named Alara Dishaean was tending her cows when she heard the braying of a mule.";
    dc:title "Michael 2-1" .

verse:Michael-2-2 a class:Verse;
    verse:bio bio:Michael_Grenadine;
    verse:concept concept:Bantu_Horse,
        concept:Curiosity;
    verse:location location:Great_Forest;
    verse:text "Curious, Alara left her farm on a Bantu Horse until she entered into the Great Forest, where she found Michael upon the forest floor, unconscious.";
    dc:title "Michael 2-2" .

Listing 6. RDF Turtle output generated from test HTML+RDFa file.

This now has broken down the RDFa markup into RDF assertions. In most cases, there was a little cheating going on - I defined objects in this example for (presumed) predefined entities. However, suppose that you only had only the text with no specific resources defined, something like:

<p><span property="verse:bioText">Michael Grenadine<span> was ...</p>

This would have generated the triple:
verse:Michael-1-1 verse:bioText "Michael Grenadine";

If you had already defined an entry for the person beforehand (<bio:Michael_Grenadine>), you could do an inference using SPARQL update:
insert {?verse verse:bio ?bio} 
where {
     ?verse verse:bioText $bioText.
     ?bio bio:fullName $bioText.
      };
Listing 7. Assigning a URI reference when given a string.

You could also create more sophisticated inferences that would take shortened forms of the person's name and the attempt to do searches on these, in the context of an already extant reference somewhere within the verse itself. I'll leave that for a later post.

Once this data is put into a triple store, it opens up some interesting possibilities. As a simple example, you could retrieve all the verse in a given book by verse in order (Listing 8).

select ?verse ?text where {
     $book book:chapter ?chapterList.
     ?chapterList rdf:rest*/rdf:first ?chapter.
     ?chapter chapter:verse ?verseList.
     ?verseList rdf:rest*/rdf:first ?verse.
     ?verse verse:text ?text.
     }
Listing 8. Retrieving verses in a book in sequential order.

where $book in this case contains the URI of the book being referenced. The rather peculiar construct of rdf:rest*/rdf:first may seem rather nonsensical, but it is an artifact of the list structure described earlier. RDF represents lists using blank-nodes, with the structure for verses in a chapter looking something like the Listing 9.

chapter:Michael-1 rdf:list   _:b0.
_:b0              rdf:first  verse:Michael-1-1;
                  rdf:rest   _:b1.
_:b1              rdf:first  verse:Michael-1-2;
                  rdf:rest   _:b2.
_:b2              rdf:first  verse:Michael-1-3;
                  rdf:rest   _:b3.
_:b3              rdf:first  verse:Michael-1-4;
                  rdf:rest   _:b4.
_:b4              rdf:first  verse:Michael-1-5;
                  rdf:rest   rdf:nil.
Listing 9. How a list is rendered as triples.

Given this, the construct rdf:rest*/rdf:first says, for a transitive relationship (rdf:rest) retrieve all rdf:rest items (including those with no rdf:rest predicate) in a path that ends with an rdf:first property. This will iterate over the list with the added benefit that the individual items do not then have to maintain their own pointers.

Why is this a benefit? In many kinds of documents, such as religious works, it's not uncommon to have "parables", which consist of one or more verses that may start in the middle of one chapter and end in another, may jump from one section to another, or my appear in a different order. than was originally given. By keeping the links external, you can add items sequentially but without being bound to needing imperative logic and exception handling. So, we can create a parable called "The parable of the mule" (Listing 10).

parable:Parable_Of_The_Mule owl:Class class:Parable;
     dc:title "The Parable of the Mule". 
     parable:verse (verse:Michael-1-2 verse:Michael-1-3 verse:Michael-1-4 verse:Michael-1-5 verse:Michael-2-1).

Listing 10. The parable of the mule.

Once you have this kind of relationship, it also becomes possible to do things like determine every parable that a given verse is used in (Listing 11), or even (assuming that verses are the only place that hold keywords) finding all parables that discuss certain concepts (Listing 12).

select ?parable where {
     ?parable parable:verse ?verseList.
     ?verseList owl:rest*/owl:first $verse.
     };
Listing 11. Retrieving all parables that have a specific $verse.

select ?parable where {
     ?parable parable:verse ?verseList.
     ?verseList owl:rest*/owl:first ?verse.
     ?verse verse:concept $concept.
     };
Listing 12. Retrieving all parables that include a specific $concept.

RDFa can appear somewhat complex at first, but the advantage to being able to encode content in RDF is that you can identify relationships between entities, can make this information aware of other information that exists within a larger dataspace, and can make that information far more malleable and while still maintaining useful context.

Kurt Cagle is an information architect and author working for Avalon Consulting LLC, specializing in NoSQL, Semantics, XML and JSON data systems. He is the author of eighteen books on XML and web technologies, including the upcoming HTML5 Graphics with SVG and CSS for O'Reilly Media. He can be reached at kurt.cagle@gmail.com.