Name
More Powerful Manipulations
We’ve just touched the tip of the iceberg for Linux text filtering. Linux has hundreds of filters that produce ever more complex manipulations of the data. But with great power comes a great learning curve, too much for a short book. Here are a few filters to get you started.
awk
awk is a pattern-matching language. It matches data by regular expression and then performs actions based on the data. Here are a few simple examples for processing a text file, myfile.
Print the second and fourth word on each line:
$ awk '{print $2, $4}' myfilePrint all lines that are shorter than 60 characters:
$ awk 'length < 60 {print}' myfilesed
Like awk, sed is a pattern-matching engine that can perform manipulations on lines of text. Its syntax is closely related to that of vim and the line editor ed. Here are some trivial examples.
Print the file with all occurrences of the string “red” changed to “hat”:
$ sed 's/red/hat/g' myfile
Print the file with the first 10 lines removed:
$ sed '1,10d' myfile
m4
m4 is a macro-processing language and command. It locates keywords within a file and substitutes values for them. For example, given this file:
$ cat myfile My name is NAME and I am AGE years old ifelse(QUOTE,yes,No matter where you go... there you are)
see what m4 does with substitutions for NAME, AGE, and QUOTE:
$ m4 -DNAME=Sandy myfile
My name is Sandy and I am AGE years old$ m4 -DNAME=Sandy -DAGE=25myfile My name is Sandy and I am 25 years old $ m4 -DNAME=Sandy -DAGE=25 -DQUOTE=yes ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access