What this error means:
You're not giving the parser the XML you think you are.
A piggy bank of commands, fixes, succinct reviews, some mini articles and technical opinions from a (mostly) Perl developer.
Jump to
Showing posts with label parser. Show all posts
Showing posts with label parser. Show all posts
Stop XML::LibXSLT and XML::LibXML connecting to w3.org
If you have this line in your XML/XSL:
<!DOCTYPE xml PUBLIC "-//W3C//DTD XHTML 1.0 Strict//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-strict.dtd">
Then your parser may connect to w3.org for validation. If w3 blocks you for hammering their server, then you may get this error message: http error : Operation in progress
1.70
$parser = XML::LibXML->new(load_ext_dtd => 0);
< 1.70
$parser = XML::LibXML->new();
$parser->load_ext_dtd(0);
<!DOCTYPE xml PUBLIC "-//W3C//DTD XHTML 1.0 Strict//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-strict.dtd">
Then your parser may connect to w3.org for validation. If w3 blocks you for hammering their server, then you may get this error message: http error : Operation in progress
http error : Operation in progress
If you don't want/need validation against the DTD (and you probably don't), you can turn off that feature in XML::LibXML:1.70
$parser = XML::LibXML->new(load_ext_dtd => 0);
< 1.70
$parser = XML::LibXML->new();
$parser->load_ext_dtd(0);
XSLT1 processing, using Perl
Example taken from XML::LibXSLT docs, code commented and replaced because of old versions of libs:
$ perl -MXML::LibXML -le'print $XML::LibXML::VERSION'
1.69
$ perl -MXML::LibXSLT -le'print $XML::LibXSLT::VERSION'
1.59
Code:
#!/usr/bin/perl
use strict;
use warnings;
use XML::LibXSLT;
use XML::LibXML;
my $xslt = XML::LibXSLT->new();
#my $source = XML::LibXML->load_xml(location => 'file.xml');
my $parser = XML::LibXML-> new();
my $source = $parser->parse_file('file.xml');
#my $style_doc = XML::LibXML->load_xml(location=>'file.xsl', no_cdata=>1);
my $parser2 = XML::LibXML-> new();
my $style_doc = $parser2->parse_file('file.xsl');
my $stylesheet = $xslt->parse_stylesheet($style_doc);
my $results = $stylesheet->transform($source);
#print $stylesheet->output_as_bytes($results);
print $results->toString;
$ perl -MXML::LibXML -le'print $XML::LibXML::VERSION'
1.69
$ perl -MXML::LibXSLT -le'print $XML::LibXSLT::VERSION'
1.59
Code:
#!/usr/bin/perl
use strict;
use warnings;
use XML::LibXSLT;
use XML::LibXML;
my $xslt = XML::LibXSLT->new();
#my $source = XML::LibXML->load_xml(location => 'file.xml');
my $parser = XML::LibXML-> new();
my $source = $parser->parse_file('file.xml');
#my $style_doc = XML::LibXML->load_xml(location=>'file.xsl', no_cdata=>1);
my $parser2 = XML::LibXML-> new();
my $style_doc = $parser2->parse_file('file.xsl');
my $stylesheet = $xslt->parse_stylesheet($style_doc);
my $results = $stylesheet->transform($source);
#print $stylesheet->output_as_bytes($results);
print $results->toString;
XML parsing fun with Perl
- XML::Simple - No XPath, just simple nested hashrefs. Good options for collapsing keys. Useful for configuration files.
- XML::Twig - Designed to handle sub-trees of very large documents, fast.
- XML::XPath - find() method produces a nodelist, which is just an arrayref (see XML::XPath::XMLParser). It's dumb to have to "re-parse" nodes in order to look further down. Maybe I'm missing something?
- XML::TreeBuilder or HTML::TreeBuilder - No XPath, but a splendid look_down() method. Returns objects on which you can also call look_down().
- XML::LibXML and XML::LibXML::XPathContext - The best.
- But HTML source often requires tidying into XML
- And there is fun to be has converting entities
use XML::LibXML;
use XML::LibXML::XPathContext;
$parser = XML::LibXML->new();
$xml = $parser->parse_string($self->raw_xml);
$xpath = XML::LibXML::XPathContext->new($xml->documentElement);
$xpath->registerNs('atom', 'http://www.w3.org/2005/Atom');
print $xpath->findvalue("/atom:feed/atom:title");
Nothing seems good for testing XML paths, tags and attributes.
Native Perl configuration file
Define a configuration file as a Perl package, containing hash data:
$ cat data
package d;
$a = {
1 => 2,
b => 'c',
};
Use this method to read the data straight into Perl:
$ perl -MData::Dumper -le'eval { require "./data"; }; print Dumper($d::a)'
$VAR1 = {
'1' => 2,
'b' => 'c'
};
This avoids the need to employ a parser, as the data is already in a Perl data structure format.
NOTE: You must specify the absolute path to the data file. I found that 'data' failed, where './data' succeeded.
If you leave out the package declaration in the data file, you can access the hash using $main::a
Subscribe to:
Posts (Atom)