A piggy bank of commands, fixes, succinct reviews, some mini articles and technical opinions from a (mostly) Perl developer.

Jump to

Quick reference

Showing posts with label files. Show all posts
Showing posts with label files. Show all posts

Windows 10 robocopy basics

robocopy c:\temp\source c:\temp\destination /E /DCOPY:DAT /R:10 /W:3

Delete all files except one on linux

Glob

  • shopt -s extglob
  • rm -v !("filename")
  • rm -v !("filename1"|"filename2")
  • rm -v !(*.zip|*.odt)
  • shopt -u extglob
Find
  • find /directory/ -type f -not -name 'PATTERN' -delete
  • find /directory/ -type f -not -name 'PATTERN' -print0 | xargs -0 -I {} rm {}
  • find /directory/ -type f -not -name 'PATTERN' -print0 | xargs -0 -I {} rm [options] {}
  • find . -type f -not \(-name '*gz' -or -name '*odt' -or -name '*.jpg' \) -delete
Glob ignore
  • cd test
  • GLOBIGNORE=*.odt:*.iso:*.txt
  • rm -v *
  • unset GLOBIGNORE

Batch deleting files in Windows

List all the jpg files:

dir /s *.jpg /b > jpgs.txt

Delete all the jpg files:

del /s /q /f /a *.jpg



Flat file vs database

Flat file
  • Easy to set up, only have to consider local file permissions
  • Easy to implement in an ad-hoc way, ideal for a prototype
    • Plain text for very simple things like a list
    • JSON/YAML/Perl for complex data structures
  • Doesn't work when the app is load balanced across multiple servers
  • Not automatically backed up
  • Amazon S3 is relatively expensive
  • Have to make your own model
Database
  • Requires up-front schema design - more work, but forces you to consider design of data 
  • You can put business logic in the ResultSet models
  • Requires an instance to be provisioned
  • Easy cross-referencing of data
  • Works when the app is load balanced across multiple servers
  • Backup-as-a-service (i.e. replication)
  • Amazon RDS is cheaper than S3
  • Get the model for free with ORM
Conclusion

For production services that have redundancy (load balanced), always use a database unless the overhead of setting one up for the first time is considered too high for the business.

How to find who/what is changing a file in linux

A list:
  • inotifywait - simple, attended
  • auditctl - powerful, old school
  • file system monitoring - more involved, more complete

See https://pinboard.in/u:random/t:audit

Create a lot of random data files on linux

# Create 10Mb random data

dd bs=1024 count=10240 < /dev/urandom > data

# Split into files of 10K each

split -a 5 -b 10240 data 'split.'

Which data file format?

List:

  • CSV
  • TSV
  • INI
  • XML
  • YAML
  • JSON
  • RDB
  • custom

How to transfer files between Windows 7 and Android 4

Settings | Storage | (three dots in top-right) | USB computer connection | Camera (PTP)

Then go to Developer Options (read elsewhere how to get this), and tick 'USB debugging'.

Voila - your device will appear in My Computer | Portable Devices.

Copy the files to a subdirectory of the Android Camera device in Windows, and move them into the right place from within Android. e.g. using ES File Explorer

Syntax highlight in vim by file extension

How to add file extension syntax highlighting to vi.
Whack this in your ~/.vimrc

au BufNewFile,BufRead *.tt set filetype=perl

(source)

STDERR and STDOUT appear in the wrong order when piping to a file

a Perl script - test.pl:

print "OUT 1\n";
print STDERR "ERR 2\n";
print "OUT 3\n";
print STDERR "ERR 4\n";
print "OUT 5\n";

Run the script, piping all output into a file:

perl test.pl &> file.log

cat file.log

ERR 2
ERR 4
OUT 1
OUT 3
OUT 5

Output is in the wrong order :(

Run it again using unbuffer:

unbuffer perl test.pl &> file.log

cat file.log

OUT 1
ERR 2
OUT 3
ERR 4
OUT 5

Output is in the correct order :)

(source)

Another way:

script -c 'perl ~/temp/stream.pl' file.log

cat file.log

(source)

How to capture linux "time" output

Pipe time output into a file like this:

{ time ls; } 2>time.output

Don't forget the semicolon!

(source)

How to sync over FTP

Don't re-invent the wheel.

1) Install rsync instead, which is designed with syncing in mind.

2) ftpsync was written a decade ago. Perhaps it has been updated since.

3) lftp syncs over FTP and is being actively maintained.

4) Perl package turbo-ftp-sync may also fit the bill.

Move directory on FTP server

without downloading and re-uploading:

rnfr source_dir
rnto target_dir

Developing over FTP

Using Eclipse and RSE plugin:
  • Logs in okay
  • But seems to always be closing directory sub-trees I've opened in 'Remote Systems' panel
  • It wants to do an 'ls' on every parent directory in the hierarchy, every time I change files. Very slow.
Using CurlFtpFS:
  • Install in Ubuntu from the software centre
    • to mount: curlftpfs -v username:password@server.example.com local_dir/
    • to unmount: fusermount -u local_dir/
  • Performance is unusably slow... Eclipse wants to know about every single file on the remote server, not just the ones I'm editing.
  • Saving the file through Eclipse using CurlFtpFS takes even longer than it would to upload it separately.
Using FileZilla and gedit:
  • Browsing is fast
  • Editing is fast
  • But you have to navigate to every file manually each session, it doesn't remember what you had open
  • The system of holding temp files locally is not ideal, but FileZilla detects changes in them very nicely
  • You still can't grep over the files as they're all held remotely, and it's a bit fiddly to navigate to a specific file
    Editing locally with Eclipse, and uploading after changes are made:
    • Write an Ant script to upload any modified files?
    • ...to investigate

    How to install perl modules into a local directory

    instead of:
    perl Makefile.PL
    try:
    perl Makefile.PL PREFIX=/path/to/your/directory/perllibs
    or:
    perl Build.PL PREFIX=/path/to/your/directory/perllibs

    Delete a file starting with a dash

    You accidentally created a file called "-.log"

    To delete it, instead of
    rm -.log
    try
    rm ./-.log

    svn: Files skipped

    Is subversion displaying a message that some files/directories were "skipped" ?
    If you are using symlinks, try cd'ing into the directory, or using the full path instead.

    Create Eclipse project with existing files

    • Have your workspace as the directory above the directory containing your files.
    • Create the project with a name that is the same as your files directory.
    The problem is, if there are lots of (unrelated) files, Eclipse will take a long time to load up.

    There are other options, i.e.:

    • Create a new folder
    • Click the 'Advanced' button
    • Link to an external folder

    sshfs on linux

    • sshfs -o uid=1000 -o gid=1000 example.com:/workplace /workplace
    Where 1000 is your user and group ID from /etc/passwd

    That's it!