Newspapers contain complex layout and content structure. While layout refers to margins, white spaces, columns, placement of headings etc, content refers to the information inside a paragraph such as type of content in the paragraph, is it enumeration, bullet lists, addresses, etc. When we automatically analyze text with complex layout and content structure such as the one we find in newspapers, information about the different elements these capture becomes very important, in particular if we want to draw objective conclusions about the evaluation results. Therefore we need to encode information about things like: how many columns are in the page, where does table of contents appear, what is the type of content in the paragraph, is it enumeration, bullet lists, addresses, etc. This kind of encoding is often referred to as structural markup.
There is no standard structural markup of newspapers. In our OCR project we had to first, before performing systematic manual analysis of the material, define a set of metadata that describes the material. For this we had to inspect the material.