Creating KVTML Files
KWordQuiz, KVocTrain, and other KDE-based programs use the KVTML file format for their data files. The format is just a fairly simple XML format but, unfortunately, there doesn't seem to be anything available to convert a text file to this format. So, once again, AWK to the rescue.
While I was extracting data from a fairly convoluted HTML file which required some more awk, the generic idea was to take some easy-to-make text format and build the needed KVTML file. I decided to use one line per record with a "|" as the field separator. There is nothing magic about using this character—you just need something that does not appear in the data itself.
My "knowledge" of the KVTML format comes from creating a file with KWordQuit and then taking a look at it. I don't really understand with rows and columns is used for but it doesn't seem important. Without further ado, here is the awk program.
# bar2kvtml.as -- convert |-separated lines into a kvtml file
# invoke as follows:
# awk -f bar2kvtml.as [l1=first_label] [l2=second_label] filename(s)
# l1 and l2 are optional column labels
BEGIN {
l1 = "Column 1" # default labels
l2 = "Column 2"
FS = "|" # field separator
print "<?xml version=\"1.0\"?>"
print "<!DOCTYPE kvtml SYSTEM \"kvoctrain.dtd\">"
print "<kvtml"
print "generator=\"bar2kvtml.as\""
print "cols=\"2\""
print "lines=\"50\""
first++
}
first { # output header with first line
nfn = FILENAME
sub(/\..*$/, ".kvtml", nfn)
print "title=\"" nfn "\">"
first = 0
print " <e>"
print " <o width=\"250\" l=\"" l1 "\">" $1 "</o>"
print " <t width=\"250\" l=\"" l2 "\">" $2 "</t>"
print " </e>"
next
}
{ # all subsequent lines
print " <e>"
print " <o>" $1 "</o>"
print " <t>" $2 "</t>"
print " </e>"
}
END {
print "</kvtml>"
}
The BEGIN block in the awk script is executed before the input file is opened. First it sets default column labels. They can be overridden on the command line. It sets the field separator (FS) to "|" and outputs most of the boilerplate. It does not, however, output the line with the filename in it as I create that filename from the input filename (replacing everything after a dot with ".kvtml". This has to be done later as FILENAME is not yet set. The variable first is set to indicate this remaining work is yet to be done.
The code executed if first is set outputs the remainder of the boilerplate plus the first record data along with the column tags—either those supplied on the command line or the defaults that were set in the BEGIN block. Finally, first is set to zero so this block will not be executed again. The next statement causes awk to skip to the next record rather than continuing processing on the current one.
For each subsequent input line, the default (no condition) block of code is executed. When the end of file is reached, the END block of code is executed.
Output is sent to standard output. You can redirect it to wherever you want. You can change the value of FS and RS to handle different imput formats. If you need to process an input format that requires a bit more processing, just add a unconditional block right after the BEGIN block that massages the data into a reasonable format. Judicious use of getline and next along with sub and gsub functions should be able to handle most anything.
Phil Hughes
Trending Topics
| You Need A Budget | Feb 10, 2012 |
| The Linux powered LAN Gaming House | Feb 08, 2012 |
| Creating a vDSO: the Colonel's Other Chicken | Feb 06, 2012 |
| Your CMS Is Not Your Web Site | Feb 01, 2012 |
| Casper, the Friendly (and Persistent) Ghost | Jan 31, 2012 |
| Razor-qt 0.4 - Qt based Desktop Environment | Jan 30, 2012 |
- Fun with ethtool
- Linux-Based X Terminals with XDMCP
- 100% disappointed with the decision to go all digital.
- Readers' Choice Awards 2011
- Parallel Programming with NVIDIA CUDA
- You Need A Budget
- Validate an E-Mail Address with PHP, the Right Way
- The Linux powered LAN Gaming House
- The Linux RAID-1, 4, 5 Code
- Python for Android
- Gnome3 is such a POS. No one
4 hours 45 min ago - Gnome 3 is the biggest POS
4 hours 55 min ago - I didn't knew this thing by
10 hours 59 min ago - Author's reply
14 hours 24 min ago - Link to modlys
15 hours 31 min ago - I use YNAB because of the
15 hours 42 min ago - Search
20 hours 45 min ago - Question
21 hours 8 min ago - for the record
21 hours 11 min ago - That's disappointing. Thanks
23 hours 34 min ago





Comments
Hi, For some reason I get an
Hi,
For some reason I get an error on the line:
sub(/\..*$/, ".kvtml", nfn)
What is that line actually for? Thanks.
- hp 6735s