A piggy bank of commands, fixes, succinct reviews, some mini articles and technical opinions from a (mostly) Perl developer.

Jump to

Quick reference

Showing posts with label parser. Show all posts
Showing posts with label parser. Show all posts

parser error : Start tag expected, '<' not found

What this error means:
You're not giving the parser the XML you think you are.

Stop XML::LibXSLT and XML::LibXML connecting to w3.org

If you have this line in your XML/XSL:

<!DOCTYPE xml PUBLIC "-//W3C//DTD XHTML 1.0 Strict//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-strict.dtd">

Then your parser may connect to w3.org for validation. If w3 blocks you for hammering their server, then you may get this error message: http error : Operation in progress

http error : Operation in progress

If you don't want/need validation against the DTD (and you probably don't), you can turn off that feature in XML::LibXML:

1.70
$parser = XML::LibXML->new(load_ext_dtd => 0);

< 1.70
$parser = XML::LibXML->new();
$parser->load_ext_dtd(0);

XSLT1 processing, using Perl

Example taken from XML::LibXSLT docs, code commented and replaced because of old versions of libs:
$ perl -MXML::LibXML -le'print $XML::LibXML::VERSION'

1.69

$ perl -MXML::LibXSLT -le'print $XML::LibXSLT::VERSION'

1.59


Code:

#!/usr/bin/perl

use strict;
use warnings;
use XML::LibXSLT;
use XML::LibXML;

my $xslt = XML::LibXSLT->new();

#my $source = XML::LibXML->load_xml(location => 'file.xml');
my $parser = XML::LibXML-> new();
my $source = $parser->parse_file('file.xml');

#my $style_doc = XML::LibXML->load_xml(location=>'file.xsl', no_cdata=>1);
my $parser2 = XML::LibXML-> new();
my $style_doc = $parser2->parse_file('file.xsl');

my $stylesheet = $xslt->parse_stylesheet($style_doc);

my $results = $stylesheet->transform($source);

#print $stylesheet->output_as_bytes($results);
print $results->toString;

XML parsing fun with Perl

  • XML::Simple - No XPath, just simple nested hashrefs. Good options for collapsing keys. Useful for configuration files.
  • XML::Twig - Designed to handle sub-trees of very large documents, fast.
  • XML::XPath - find() method produces a nodelist, which is just an arrayref (see XML::XPath::XMLParser). It's dumb to have to "re-parse" nodes in order to look further down. Maybe I'm missing something?
  • XML::TreeBuilder or HTML::TreeBuilder - No XPath, but a splendid look_down() method. Returns objects on which you can also call look_down().
  • XML::LibXML and XML::LibXML::XPathContext - The best.
    • But HTML source often requires tidying into XML
    • And there is fun to be has converting entities
use XML::LibXML;
use XML::LibXML::XPathContext;
$parser = XML::LibXML->new();
$xml = $parser->parse_string($self->raw_xml);
$xpath = XML::LibXML::XPathContext->new($xml->documentElement);
$xpath->registerNs('atom', 'http://www.w3.org/2005/Atom');
print $xpath->findvalue("/atom:feed/atom:title");


Nothing seems good for testing XML paths, tags and attributes.

Native Perl configuration file

Define a configuration file as a Perl package, containing hash data:

$ cat data
package d;

$a = {
    1 => 2,
    b => 'c',
};

Use this method to read the data straight into Perl:

$ perl -MData::Dumper -le'eval { require "./data"; }; print Dumper($d::a)'
$VAR1 = {
          '1' => 2,
          'b' => 'c'
        };

This avoids the need to employ a parser, as the data is already in a Perl data structure format.

NOTE: You must specify the absolute path to the data file. I found that 'data' failed, where './data' succeeded.

If you leave out the package declaration in the data file, you can access the hash using $main::a