This show has been flagged as Clean by the host.
ether
This series is dedicated to exploring little-known—and occasionally useful—trinkets lurking in the dusty corners of UNIX-like operating systems.
I frequently find myself reaching for the
cut
utility when writing scripts to extract one piece of data from a line, or to select specific fields from a log file. While I am familiar with its counterpart,
paste
, I don't employ it very often because I don't typically need its functionality.
This perhaps has to do with the fact that I rarely work with text files containing lists. For shorter lists, I usually end up using a spreadsheet and for larger ones, a relational database. Both are valuable tools with their own strengths and weaknesses, but it is good to also know about standard utilities for working with lists. After uploading UNIX Curio #8 (
1
, but can explain how it works. Briefly, it is a rough opposite of
cut
—when given multiple files as arguments, it assembles the first line from each one separated by tabs, then the second line, and so on. Instead of tabs, a different delimiter can be chosen with the
-d
option. Another option is
-s
, which swaps rows and columns so that the contents of each named file would appear on one line. While
paste
itself doesn't qualify as a UNIX Curio in my opinion, there is one feature that does: a hyphen can be given as an argument multiple times. In this special case, the output is taken line by line from standard input, but is spread across as many columns as there are hyphens.
Example of using
paste
to turn the output of
ls
into columns. Because these columns are separated by tabs, they don't necessarily line up when a filename is eight or more characters long. The
-1
is not required for the second
ls
command since that behavior is implied when output isn't going to a terminal. The
-C
option to
ls
usually gives nicer-looking output on a terminal—also, it lists in ascending order down by column. (Most implementations default to
-C
when output goes to a terminal.) If you want items ascending along rows like the
paste
example does, try
ls -x
instead.
$ ls -1 /proc/net
anycast6
arp
bnep
connector
dev
dev_mcast
dev_snmp6
fib_trie
fib_triestat
hci
icmp
icmp6
if_inet6
igmp
igmp6
ip6_flowlabel
ip6_mr_cache
ip6_mr_vif
ip_mr_cache
ip_mr_vif
ip_tables_matches
ip_tables_names
[...35 more entries not shown...]
$ ls /proc/net | paste - - - -
anycast6 arp bnep connector
dev dev_mcast dev_snmp6 fib_trie
fib_triestat hci icmp icmp6
if_inet6 igmp igmp6 ip6_flowlabel
ip6_mr_cache ip6_mr_vif ip_mr_cache ip_mr_vif
ip_tables_matches ip_tables_names ip_tables_targets ipv6_route
l2cap mcfilter mcfilter6 netfilter
netlink netstat packet protocols
psched ptype raw raw6
rfcomm route rt6_stats rt_acct
rt_cache sco snmp snmp6
sockstat sockstat6 softnet_stat stat
tcp tcp6 udp udp6
udplite udplite6 unix wireless
xfrm_stat
$ ls -C /proc/net
anycast6 if_inet6 l2cap rfcomm tcp
arp igmp mcfilter route tcp6
bnep igmp6 mcfilter6 rt6_stats udp
connector ip6_flowlabel netfilter rt_acct udp6
dev ip6_mr_cache netlink rt_cache udplite
dev_mcast ip6_mr_vif netstat sco udplite6
dev_snmp6 ip_mr_cache packet snmp unix
fib_trie ip_mr_vif protocols snmp6 wireless
fib_triestat ip_tables_matches psched sockstat xfrm_stat
hci ip_tables_names ptype sockstat6
icmp ip_tables_targets raw softnet_stat
icmp6 ipv6_route raw6 stat
$ ls -x /proc/net
anycast6 arp bnep connector dev
dev_mcast dev_snmp6 fib_trie fib_triestat hci
icmp icmp6 if_inet6 igmp igmp6
ip6_flowlabel ip6_mr_cache ip6_mr_vif ip_mr_cache ip_mr_vif
ip_tables_matches ip_tables_names ip_tables_targets ipv6_route l2cap
mcfilter mcfilter6 netfilter netlink netstat
packet protocols psched ptype raw
raw6 rfcomm route rt6_stats rt_acct
rt_cache sco snmp snmp6 sockstat
sockstat6 softnet_stat stat tcp tcp6
udp udp6 udplite udplite6 unix
wireless xfrm_stat
The
paste
command has limitations—the files you give it must all be already arranged in the same order, and if any file is missing a value, it must have a blank line so that subsequent lines will match up correctly. The files do
not
necessarily have to be sorted alphabetically, but whatever order they are in has to be the same. Check out HPR episodes
for some more background on the
paste
utility.
Example of using
paste
with files where some values are empty. Bob works from home so doesn't have an office assigned, and the laboratory Carol works in doesn't have a phone. This relies on the fact that the same line number in every file relates to the same person/entry.
$ cat names
Alice
Bob
Carol
Dave
$ cat offices
203
Lab6A
117
$ cat phones
+1 212-555-1234
+1 919-555-2345
+1 212-555-1278
$ paste names offices phones
Alice 203 +1 212-555-1234
Bob +1 919-555-2345
Carol Lab6A
Dave 117 +1 212-555-1278
Our second UNIX Curio for today is
2
, which has a bit more sophistication. It operates on two files, which can have multiple columns, and combines them using the join field. By default, the first column/field in each file is the join field, and only entries that exist in both files are printed. The
-1
and
-2
options can be used to join on a different field, and
-o
selects specific fields to be output. To make it so lines with missing entries also appear, you need to use the
-a
option, but an actual empty string with separator won't be printed unless
-o
is also present and includes the field.
The default field separator character is one or more "blanks" in the current locale—for the POSIX locale, this means a space or a horizontal tab. The
-t
option selects a different character and also removes the treatment of multiple occurrences as a single separator, making it possible to have an empty field in one or both of the files. By default, a single space is used to separate fields in the output. If
-t
is given, the same character is used for separating fields in both input and output. You would need to pipe output through another tool like
tr
if you wanted to have a different separator in the output.
The
join
utility might be an improvement over
paste
in some cases, since the join field makes it a little easier to identify which entries match up across files. It is limited to operating only on two files (one of which can be standard input), so combining more than that requires either creating temporary intermediate files or chaining together
join
commands in a pipeline. Another requirement is that all files must already be sorted in the current locale.
Example showing how
join
can be used with two tab-separated lists. The LC_ALL assignment forces
join
to sort using the C (POSIX) locale instead of whatever might be set in your environment. The "@" on the header line has no special meaning; it is just there to make sure it sorts before any letters or numbers (in the C locale; it might not in other locales). Note that if
-t
were not specified,
plist
would be treated as having three fields because of the space separating the country code from the rest of the phone number.
$ export tab="$(printf '\t')" #To more easily use tab characters below
$ cat olist
@Name Office
Alice 203
Carol Lab6A
Dave 117
$ cat plist
@Name Phone
Alice +1 212-555-1234
Bob +1 919-555-2345
Dave +1 212-555-1278
$ LC_ALL=C join -t "$tab" olist plist
@Name Office Phone
Alice 203 +1 212-555-1234
Dave 117 +1 212-555-1278
$ LC_ALL=C join -t "$tab" -a 1 -a 2 olist plist
@Name Office Phone
Alice 203 +1 212-555-1234
Bob +1 919-555-2345
Carol Lab6A
Dave 117 +1 212-555-1278
$ #By default, join acts as if empty fields don't exist; use -o to include
$ LC_ALL=C join -t "$tab" -a 1 -a 2 -o 0,1.2,2.2 olist plist
@Name Office Phone
Alice 203 +1 212-555-1234
Bob +1 919-555-2345
Carol Lab6A
Dave 117 +1 212-555-1278
$ #The -e option sets a placeholder to use for empty fields
$ LC_ALL=C join -t "$tab" -e "(none)" -a 1 -a 2 -o 0,1.2,2.2 olist plist
@Name Office Phone
Alice 203 +1 212-555-1234
Bob (none) +1 919-555-2345
Carol Lab6A (none)
Dave 117 +1 212-555-1278
The brief description for
join
is "relational database operator"—I won't dispute that, but in my view it offers far fewer capabilities than people would expect from today's relational databases. I would imagine that when most people think of those they have Structured Query Language (SQL) in mind, which offers a lot more flexibility and functions to operate on data. However, I can see how
join
could be suitable for simple operations.
Our last UNIX Curio for today relates to
4
in 1973. What
did
come as a shock to me is that both
cut
and
5
, and were actually preceded by
6
from 1979. I assumed that at least
cut
would have been around far earlier, given its usefulness and how firmly established it is, but I suppose it just
seems
to have been with us forever.
As mentioned, I don't typically manage data as text files containing lists, and I probably won't start using the
join
utility or these features of
paste
and
sort
very much. But it is still useful to know that they exist and how they work. Hopefully this episode has taught you a bit about them.
References:
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/join.html
https://archive.org/details/a_research_unix_reader/page/n19/mode/1up
https://man.cat-v.org/unix_7th/1/join
Provide feedback on this episode.
SOCIAL SHARE CARD GENERATOR