A lot of the time it seems that Perl programmers write a regex like
/^foo$/
(where “foo” is some arbitrary regex pattern)
but from the context it seems like the intention of the programmer was
to make it so that the string to be searched must match the pattern
“foo”, and there must be nothing between the beginning of “foo” and the
beginning of the string to be searched, and there must be nothing
between the end of “foo” and the end of the string to be searched.
But of course the regex doesn’t do that.
The metacharacter ‘$’ matches not only at the end of the string to be
searched but also just before a newline character at the end of the
string to be searched. (Of course when the ‘m’ flag is specified, ‘$’
behaves differently. But I’d like to concentrate on the behaviour
without the ‘m’ flag for the time being.)
So the above pattern will match “foo” and “foo\n”.
Is that what the programmer really wanted? I think in many cases not.
So how can we make the pattern match exactly at the end of the string to
be searched?
The answer is to use the metacharacter ‘\z’. This matches exactly at the
end of the string to be searched.
So to make a regex that matches the pattern “foo”, and with the
beginning and end of the pattern bound to the end of the string to be
searched, we could write:
/^foo\z/
=====
Here endeth the bit about doing the minimum to make the code correctly
reflect the intention of the programmer. The rest is about style,
personal preference, readability, etc.
=====
Some might say that using ‘^’ to match the beginning of the string to be
searched and ‘\z’ at the end is a bit dicey because the meaning of ‘^’
is changed if the ‘m’ flag is used but the meaning of ‘\z’ isn’t. It
would be nice if there was a metacharacter which exactly matched the
beginning of the string to be searched, regardless of the ‘m’ flag.
Fortunately there is, ‘\A’. Using that would give:
/\Afoo\z/
But because ‘\A’ ends with a letter, the regex can be a bit hard to
parse if the ‘\A’ is followed by a pattern which begins with a letter.
So some might say that it might be a good idea to use the ‘x’ flag to
allow whitespace inside the regex. That would give
/ \A foo \z /x
Though I find it a bit hard to read when slash delimiters are combined
with a few initial-backslash metacharacters, so I prefer to use a
different delimiter. So I would think to use something like:
m{ \A foo \z }x
- by Bill Blunn
See also http://perldoc.perl.org/perlre.html#Regular-Expressions
A piggy bank of commands, fixes, succinct reviews, some mini articles and technical opinions from a (mostly) Perl developer.
Jump to
Showing posts with label regex. Show all posts
Showing posts with label regex. Show all posts
Bash parameter expansion
Have you ever seen ## (hash hash/pound pound) or %% (percent percent) inside a bash script and wondered what it means? It's a form of parameter expansion that allows you to manipulate your strings using regexes.
All the following commands work with a variable called $string.
'pattern' means bash pattern, or can also be an ordinary string:
All the following commands work with a variable called $string.
'pattern' means bash pattern, or can also be an ordinary string:
- String length: ${#string}
- Extract a substring: ${string:position}
- Extract a substring, specifying length: ${string:position:length}
- Delete shortest match of substring from the beginning of string: ${string#substring}
- Delete shortest match of substring from the end of string: ${string%substring}
- Delete longest match of substring from the beginning of string: ${string##substring}
- Delete longest match of substring from the end of string: ${string%%substring}
- Find and replace first substring: ${string/pattern/replacement}
- Find and replace all substrings: ${string//pattern/replacement}
Thanks to TheGeekStuff
Labels:
bash,
linux,
pattern matching,
regex,
strings
Combine multiple lines in multiple files into a single spreadsheet row
for k in $(ls *.xml); do echo "$(perl -e'print "$ARGV[0]\t"' $k)""$(head -14 $k | egrep 'h:shortdesc|dc:title' | perl -lne'($t)=$_=~m{([^<]+)} if ! $t; ($d)=$_=~m{([^<]+)<} if ! $d;END{print "$d\t$t"}')"; done
PHP regular expressions
/(.*)something/s
The s at the end causes the dot to match all characters including newlines.
Use Perl-like regular expression in vi
You have to escape some special characters and not others:
Escape special pluses, capturing parentheses and literal $
use \1 for the captured strings.
Regular expressions in vi
Keep the first character of a line, but then insert an 'A':
s/^\(.\)/\1A/
NOTE: parentheses and the plus sign must be escaped.
Labels:
productivity,
regex,
text editor,
vi
Regular expressions in C: Advice
If you can avoid using regular expressions in C, then *do not use them!!*
Instead try the strstr (find substring) and strtok (like perl's split) routines.
Labels:
c,
idea,
programming,
regex,
strings
Regular expressions in C
Using glibc's regex.h
- Make sure the number of expected matches (nmatch) is high enough, or there will be garbage at the end of the matchptr array
- If regexec() succeeds [returns zero], matchptr contains:
- 0: the whole regex
- 1: first parenthesis match
- if regexec() fails [returns a value], don't even look in matchptr
- Characters that must be escaped:
- parentheses
- plus signs
- If you're matching the end of line, it must be inside any parenthesis!
Labels:
c,
programming,
regex
XSL regular expressions
NOTE: Variables must be declared in a root <xsl:template> node.
<!-- extract the ID -->
<xsl:variable name="id">
<xsl:analyze-string select="doc:entry/doc:id" regex="id/(\d+)">
<xsl:matching-substring>
<xsl:value-of select="regex-group(1)"/>
</xsl:matching-substring>
</xsl:analyze-string>
</xsl:variable>
<!-- extract the ID -->
<xsl:variable name="id">
<xsl:analyze-string select="doc:entry/doc:id" regex="id/(\d+)">
<xsl:matching-substring>
<xsl:value-of select="regex-group(1)"/>
</xsl:matching-substring>
</xsl:analyze-string>
</xsl:variable>
Subscribe to:
Posts (Atom)