Showing posts with label XML. Show all posts
Showing posts with label XML. Show all posts

Tuesday, June 14, 2011

An XML Primer

Wow. I just did a "Bing" on "XML" and found 88,300,000 results. The third facet on the results page (with faceted search being the reason I prefer Bing) was "XML Definition." Nineteen million pages fell under the "related searches" facet of XML definition. I zapped off a few other searches of popular tech terms and three-letter acronyms (RSS, IP address, namespace, API, RDF -- and none of them had a facet called "Definition."

So what can be made of this? If you attend any digital media seminar, workshop or webinar or sit in on any content strategy, XML is de rigueur, but could be it be that people are throwing out this TLA without really knowing from whence they speak? The answer is absolutely. And solution providers, technologists and product makers are guilty of not recognizing that the community is struggling to keep up.

Cathy Palmer and I partook in a web series put on by the IDEAlliance this morning on making the case for XML. IDEAlliance is a non-profit that develops standards and best practices surrounding publishing and technology -- it offers events virtually every week of the year depending upon practice area. Cathy is a trainer from New Horizons, a nationwide IT training company. A couple hours later I was listening to a webinar by Publishing Executive -- featuring two book publishing executives. Peppered liberally throughout both webcasts was our little friend XML. And then came the question asked in a variety of ways: "But what if we don't have XML, what do we do?" Cathy did a super job explaining how you can extrapolate XML from InDesign files, while I offered that another way is to use combinations of machines (semantic analysis engines) and man (offshore) to create XML.

But how do executives create a content strategy -- determining man, machine and markup if they don't have a rudimentary understanding of what this eXtensible Markup Language is all about? The definition is easy -- the why is more complex. XML is a decade-old method of mark-up that can be used to classify and add meaning to content so that it can be organized, “sliced and diced” and repurposed.

XML tags look similar to HTML (HyperText Mark-up Language) ones, in that they both use start and end tags but that’s about it with the similarities. HTML includes a set of pre-defined formats that impact how information is rendered, eg. the command (along with it's close command ) makes a word bold. Unlike HTML, XML does not have predefined formats (although it does use the same syntax) and display commands; instead XML provides a structure so you can effectively find information again.

This format agnostic markup language means you can categorize sections of content -- find them again -- and then transform them (using style sheets) to be ready for virtually any digital channel. XML allows bodies of content to be broken down into reusable components -- for instance, maybe you would like to markup statistics within a text -- particularly if you know that you will be researching for that same type of statistic again. Or maybe you want to markup quotations by luminaries; charts by researchers, lyrics to songs, ingredients to recipes. Having the ability to search, find and reassemble these components of content is the secret to repurposing.

There's more to XML than that -- but that's the high-level basics. The key to managing is understanding what you know -- and don't know -- and filling in the gaps.

Friday, December 17, 2010

Let's Start at the ... Very End!

Apps, apps, apps. I need a good appetizer recipe for a holiday party, so I will be turning to Epicurious, AllRecipies -- and maybe even download the new Mario Batali Cooks! to find something yummy.

Why am I telling you this? Because if we want to join into the app craze -- we have to think a bit differently - like at the end, first. Or more specifically, what we want the end product to be. I know that probably sounds quite intuitive but for many companies just getting a mobile app out there seems to be more important than getting the right mobile app out there. And no wonder. With Chris Anderson and others now running around screeding the "Web is Dead" we are in a panic to push our sites to mobile.

Which is all well and good from a replication standpoint -- but not so good if what you really want is a killer app like those Angry Birds. Admittedly most of us aren't in the gaming business, but if your business is communicating with you audience -- then you need to put yourself in the shoes of your audience -- where ever those shoes may be.

For instance, the American Institute of Physics provides a digital platform for its 180 scholarly & trade journal constituents. Long journal articles are not going to do it for audiences looking for information  -- while say, atop a ladder. Instead they are building out specific apps that let readers ask the app specific questions -- what part do I need? -- and get a specific answer.

This requires a completely different mindset; instead of pushing information to the reader -- the reader is pulling the information required. Instead of a product -- media companies are providing a service. Its only by anticipating the types of questions readers might ask (under various conditions -- do they travel, are they outdoors in a field, at a desk) that publishers can begin to start developing highly sought-after applications.


A top 10 app for the Android right now is Greg Milette's Digital Sidekick. Responding to a Google challenge to develop for 'Droid, Milette chose to take advantage of the voice recognition tools -- and his love for cooking. Understanding the workflow of preparing recipes, he opted for a recipe reader that would allow the cook to ask basic questions without having to quick looking at a recipe: What temperature to preheat the oven? How much flour do I need? The app keeps track of which portion of the instruction has been read -- and can pick up where the person left off - despite any interruptions.

Milette designed hooks to import recipes from AllRecipes.com – but readers can also cut and paste from any site to put their own favorites to create their own cookbook.

Now Milette's app is unique in that he isn't packaging up his own content for an app -- he is relying on content created by others. Had he had his own repository of recipes he would have needed to prepare his content ecosystem so it would allow this re-assembly of recipes into this new talking cookbook.

In most cases, this is not an easy task - unless you are storing the content as XML -- and have an easy way to index it, find it and deliver it. Most media and enterprises have relied on databases that were built to efficiently handle structured data that is in columns and rows. Unfortunately these RDBMS databases do not handle unstructured content -- like articles, recipes, video and images -- very well at all. A flexible  content ecosystem today requires:

  • a database that is purposely built for unstructured content - perfect for XML
  • a means of enriching the content (with both semantic and administrative metadata) 
  • a way to transform the XML from one schema to another so it can be delivered

By knowing what our finished product should be, we can take an audit of our content -- and see what is missing, determine the granularity of the enrichment (should people be able to search by types of cuisine and whether or not it is an appetizer or a dessert?) and which types of additional content may be needed -- that would be brought into the ecosystem -- and re-assembled and delivered.

These are not easy tasks without the right tools, particularly if you want users to add their own -- or 3rd party content -- which may not follow the same XML schema as your recipes did.

At Intelligent Content 2011, my MarkLogic colleague Fernando Mesa and I will be giving a plenary talk on creating this very agile content ecosystem -- and offer some terrific real world examples of how others are quickly creating applications that merge in disparate types of content -- and preparing these apps for all sorts of delivery mechanisms. It should be a terrific session -- not to mention warm -- since it is in Palm Springs. You should definitely come -- and bring your ideas which will allow us to brainstorm apps that will be killer for your audience.

Sunday, August 22, 2010

6 Things About Unstructured Content You Need to Know

Unstructured data. As a writer I hate that term. I remember the first time I heard reference to it: sitting in a meeting and technical people were talking about all the unstructured content that publishers produce. How could they be speaking about articles as unstructured? If anyone has made it through kindergarten they have learned that they must follow a linguistic pattern or structure in order to communicate effectively. But in the lexicon of geekdom, any article, picture, powerpoint, video, song, user-generated content -- your kids' text messages -- are all: unstructured. Any "data" that does not fit nicely into a column or a row -- is considered unstructured.

Now, as much as I hate the term, there is a logic that places all content that didn't fall neatly into a table to be called unstructured versus structured. It was evident when classifieds first went online. Dumped from mainframes where customers paid by the character, people created their own short-hand to say 4BR House 4 Sale. Turns out though, that all that unstructure (which I prefer to say as free-form)  makes it very hard to search on. Don't believe me? Go to Craig's List -- which tries to impose structure on advertisements by putting it under broad categories and locations. Other than that -- it is pretty freestyle. Nannies are caregivers are sitters (baby or otherwise) -- and plural or otherwise. And search on one of those terms at the peril of not finding it under the other.

On the flip side are the sites that allow only structure: information fits neatly in a row or a table and has descriptive titles like Type of residence, # of Bedrooms, Siding, Price, MLS # etc. Having that structure makes it easy to query or search on that information. You probably have heard of SQL -- which is the acronym for Structured Query Language. Relational databases use SQL to find content -- by looking in the appropriate fields. Which is all well and good for content like financial information, inventories, human resources stuff. But what about the rest of the content that floats around a corporation? The memos, sales presentations, business plans, schematics, Web sites -- the stuff we sometimes refer to as: Knowledge - and which in geekdom is called unstructured content, are the digital assets we need to carefully manage.

So here's 6 Things You Need to Know About Unstructured Content.
  1. It's everywhere. Analysts, pundits and people in the know estimate that more than 80% of content produced in an enterprise is unstructured
  2. Content is containerized. Unstructured content resides in containers like .doc, .ppt, .tiff -- and you must have the right software application to read or edit it.
  3. Managing unstructured content is hard. Because content resides in containers it is hard to know what it is in each one. 
  4. XML is crucial for reuse and sharing. Sometimes called atomic or neutral format, XML is a language used to transmit content -- without burden of the container. Neutral content can be then "poured" into any template (Word, Web, PDF -- mobile apps!!!) for easier repurposing. If you have unstructured content (and most likely you have lots of it) it should be stored in an XML format.
  5. Good metadata is essential. Once content is in an XML format, enrich it with semantic metadata. This is critical to letting knowledge workers find out what each asset is about
  6. Native XML databases provide agility & efficiencies. Relational databases (RDBMS) are great for organizing and querying structured data -- while XML databases rock for unstructured content. You can make an RDBMS work with XML, but you will lose a lot in database performance (upwards of 30% is estimated by Forrester). ) -- Heck I will use a knife to tighten a screw -- but sometimes I need to go and get the Philips head.
Now let's look briefly at what all this means. You have tons of content that doesn't fit into tables and rows. Storing content in their original containers of powerpoints, word docs and PDFs makes it extraordinarily hard to share and repurpose this content -- since it's hard to search it -- and cutting and pasting becomes the only alternative. And when we are talking about sharing and repurposing, remember all the great mash-up apps that your content might be perfect for -- if only it were in a neutral format.

Look, I don't like the term Unstructured Content -- but the acronym of CTDFWICAR (Content that doesn't fit well in columns and rows) is hardly memorable -- and way too long. Semi-structured content (because XML actually follows schemas -- which makes it semi-structured, but that's for another day) is only half as horrible as unstructured. In any event, while it may be free-form -- this type of content is the lifeblood of any organization -- and deserves its own special database to help keep it valuable.

Monday, August 2, 2010

Mixing, Mashing XML into Content Derivatives

After much deliberation, I have changed jobs, moving from Nstein (now OpenText) to MarkLogic, a provider of “purpose-built databases for unstructured content,” which means we handle all that data that doesn’t fit nicely into rows and columns – you know content like documents, articles, books, graphics -- which is often (and best) represented in XML. Over the years I have written about the importance of semantic metadata but it is but a scintilla of types of metadata that can be appended to content -- as long as that all-important infrastructure is in place. It was so last year to be absorbed in knowing what content you created -- now what is important is to complement that content with information from other resources.

What does this mean to information providers, publishers and other types of media (and if you read my blogs, you know that I believe all of us are publishers!)? Well consider the iPad. Selling at a head-shaking rate of one every three seconds (despite the recession), more than 13 million will be in consumers' hands by Christmas. Add to it the hordes of other mobile devices: iPhones, Blackberries, Androids, eReaders ... and you have 12 percent of the market looking for content for their gizmos. Which in itself can be a challenge -- since content doesn't just magically play nicely on every device. Most of that content will need to flow into its own native application to really exploit the devices' features, which mean that content needs to first be available in a neutral format.

And while you are exploiting the gizmos features ... remember that if gizmos are everywhere their owners are -- there will be an increased desire for what is known in the military as situational awareness ; deriving additional, contextual information that relates to the user -- usually time (temporal data) and space (geo mapping). Think of a soldier needing to know what threats are in the area where he currently is -- or where he is going. Publishers too should think in terms of creating new derivative, situational content. For example, what types of information might a business person want while on Maple and Elm at 8am in the summer -- versus at 8pm in the winter? Or what location-based weather patterns does a commodities trader want when looking at crop futures?

This need for situational awareness provides a great opportunity for publishers to take their knowledge bases and mix it with external resources -- such as public information from NOAA, Google Maps, LinkedIn, or proprietary information from partners. The key to mixing and mashing is having content in a mutable format -- and a database that can handle it. Extensible Markup Language, or XML, is a highly flexible text format, a W3C standard that is sometimes called atomic or a neutral format. It is designed to be easily stored and retrieved -- void of any display format. Which means it "pours" nicely into any layout. MarkLogic's database in my mind then is akin to a gourmet mixing bowl that takes in XML and allows it to be stored and retrieved into any application.

XML is hardly new as it was designed for large-scale publishing, although there is an increased awareness around it due to the Web and blog feeds, and is a great way to describe unstructured data. Unstructured data can reside in regular relational databases (RDBMS) -- but they tend to get bogged down. By storing this unstructured data on a database built specifically to handle these datatypes -- you can search and retrieve much more quickly. Forrester Analyst Noel Yuhanna told me estimated that by unburdening RDBMS of unstructured data -- they saw a 30% lift in database performance, which is huge.

In any event, the real advantage to having content in XML -- and residing in a database built to handle it -- is that you can easily mix and mash it up into new types of content, ready it for new delivery platforms, or ease syndication. And you can do all of this in a matter of weeks not months -- which is terrific since we don't yet know what other new gizmo might be in readers' hands -- least of all by Christmas.