Wednesday, December 15, 2010

Delete executables and object (.o) files

Following two commands can delete all ELF executables and ELF object (.o) files.

find ./ -executable -type f|xargs -I{} file {} | grep ELF|cut -d ':' -f 1|xargs -I{} rm {}   

find ./ -name "*.o" -type f|xargs -I{} file {} | grep ELF|cut -d ':' -f 1|xargs -I{} rm {}

Some executables may not be found because its "x" bit is not set. Use following command to find them

find ./ -type f|xargs -I{} file {} | grep ELF|cut -d ':' -f 1|xargs -I{} rm {}

Note:
    -executable is not supported in old version of find.

Monday, December 13, 2010

How to change hostname in Ubuntu

Temporary change

    hostname <new_host_name>

Permanent change

  1. Edit /etc/hostname to specify your new hostname
    sudoedit /etc/hostname
  2. sudo service hostname start

Ubuntu init scripts and upstart jobs

Init script

Those scripts are located in directory /etc/init.d. Note: some of them have been converted to upstart jobs (see next section) and they should not be invoked directly.  To check whether it's a upstart job, check directory whether file /etc/init/<job>.conf exists for a specific job.

upstart jobs

http://upstart.ubuntu.com/  Use "man 5 init" to see the syntax of the conf file.

Each upstart job has a conf file in directory /etc/init/<job>.conf. You should not directly invoke the init script to start/stop the job. You should use commands initctl to do that. E.g. initctl restart hostname. 
initctl list will list upstart jobs that are running.

service

It can be used to interact with the init scripts, no matter they are upstart jobs or regular init script. For upstart jobs, it does not run /etc/init.d/<job> . Instead it runs "start <job>" directly.

E.g.   service hostname status

invoke-rc.d

Another tool to start/stop init jobs. In my opinion you should use command service because invoke-rc.d does NOT detect whether the job is a upstart job or regular init job. Usually, this is not a big deal because upstart job shell script automatically calls initctrl related commands (start, stop, reload, etc) .

Example - network

I will give an example about how network interfaces are managed by init daemon.

As you may know, ifup and ifdown can be used to bring up or down network interfaces.

/etc/network/interfaces are used by ifup and ifdown to know how you want your system to connect to the network.

Sample file

# interfaces lo and eth0 should be started when ifup -a is invoked.
auto lo eth0
# eth1 is allowed to be brought up by subsystem hotplug.
allow-hotplug eth1
# For interface lo, it should use internet protocol and it is a loopback device.
iface lo inet loopback
# Interface eth1 uses internet protocol and dhcp for configuration
iface eth1 inet dhcp

  • For upstart job networking, its config file is /etc/init/networking.conf:

description "configure virtual network devices"

start on (local-filesystems
      and stopped udevtrigger)

task

pre-start exec mkdir -p /var/run/network

exec ifup -a

Notice the last line? Yes, it invoke ifup to bring up those interfaces that are marked as "auto" in file /etc/network/interfaces.

  • Job network-interface is used when a network interface is added or removed. Its config file is /etc/init/network-interface.conf

description "configure network device"

start on net-device-added
stop on net-device-removed INTERFACE=$INTERFACE

instance $INTERFACE

pre-start script
    if [ "$INTERFACE" = lo ]; then
    # bring this up even if /etc/network/interfaces is broken
    ifconfig lo 127.0.0.1 up || true
    initctl emit -n net-device-up \
        IFACE=lo LOGICAL=lo ADDRFAM=inet METHOD=loopback || true
    fi
    mkdir -p /var/run/network
    exec ifup --allow auto $INTERFACE
end script

post-stop exec ifdown --allow auto $INTERFACE

Line "exec ifup --allow auto $INTERFACE" bring up the newly added interface if it is set to be brought up automatically.  The trigger event is "net-device-added" or "net-device-removed" which is sent by upstart-udev-bridge. It basically forwards events received from udev to init daemon. When your network interface (e.g. eth0) is detected by udev, finally a net-device-added event is sent to network-interface upstart job which runs ifup to bring it up.

  • upstart job hostname. (/etc/init/hostname.conf). It includes following line:
      exec hostname -b -F /etc/hostname
    Now you should know how to change hostname.
    Edit file /etc/hostname, run command "sudo service start hostname", "sudo start hostname", or "sudo initctl start hostname".

Recover corrupted partition table

Partition table of my linux drive was corrupted recently.  I could not start up my Ubuntu.

I burned a Ubuntu CD. But when I tried to boot into the liveCD, it always gave me errors. It seems to be a CD burning/CD drive problem. Then I made a live USB drive which worked great. Following two tools can be used to "guess" the partition table.

  • gpart
    This program is kind of old and is not maintained any longer. It can just recognizes some file systems (ext3, ext4, etc are not recognized correctly)
  • testdisk
    This is a great tool which a text UI. You can find information here. You just follow the instructions. Check the "guessed" partition table match your real partition table(if you have backup, you are lucky.).

Run fsck to check integrity of your file system.

Afterthoughts:

  1. Back up your partition table!
  2. USB drive is more stable than CD in this case.

Sunday, November 28, 2010

printf cheatsheet

Format string

  % [flag]* [minimum_field_width] [precision] [length_modifier] <conversion_specifier>

Flags

Flag Description  
# The value should be converted to an "alternate form".  
0 zero padded. By default, blank padded  
- left adjusted  
' ' (A single space) A blank should be left before a positive number (or empty  string)  produced by a signed conversion.  
+ A sign is always placed before a number produced by a signed conversion.  

Minimum Field Width

"a decimal digit string (with  non-zero first digit). If the value has fewer characters, it will be padded.  In no case does a nonexistent or small field width cause truncation of a field."

Precision

a period ('.')  followed by an optional decimal digit string. It has different meanings for different conversions.

Precision Description  
d, i, o, u, x, and X minimum number of digits to appear printf("%.2d", 1) ==> "01"
a, A, e, E, f, and F number of digits to appear after  the radix  character printf("%.2f",0.1) ==> "0.10"
g and G maximum number of significant digits  
s and S maximum number of characters to be printed from a string printf("%.2s","hello") ==> "he"

Length Modifier

For each conversion specifier, there is expected argument type. For example, for conversion d,  and i, type of arguments should be int. Length Modifiers can be used to specify argument types rather than expected type by default.

Modifier Conversion Argument types
hh d, i, o, u, x, or X signed char or unsigned char
n signed char*
h d, i, o, u, x, or X short int or unsigned short int
b short int*
l d, i, o, u, x, or X long int, or unsigned long int
n long int*
c wint_t
s wchar_t
ll d, i, o, u, x, or X long long int, or unsigned long long int
n long long int*
L a, A, e, E, f, F, g, or G long double
j d, i, o, u, x, or X intmax_t, uintmax_t
z d, i, o, u, x, or X size_t, ssize_t
t d, i, o, u, x, or X ptrdiff_t

Conversion specifier

conversion arguments notation  
d, i int argument signed decimal notation  
o, u, x, X unsigned int unsigned octal, unsigned decimal, unsigned hexdec notation abcedf are used for x.
ABCDEF are used for X.
e, E double rounded and converted in the style [-]d.ddde±dd precision is 6 by default.
f, F double rounded  and  converted  to  decimal  notation in the style [-]ddd.ddd precision is 6 by default.
g, G double converted in style f or e (or F or E for G  conversions).
Style e  is  used  if the  exponent from its conversion is less than -4 or greater than or equal to the precision.
 
a, A double converted to hexadecimal notation.
for a, (using the letters abcdef) in the style [-]0xh.hhhhp±d;
C99;  not in SUSv2
c int converted to an unsigned char, and the resulting character is written  
s const char* Characters from the array are written.  
p void * printed in hexadecimal (as if by %#x or %#lx)  
n int * The number of characters written so far is stored into the integer  indicated  by  the int * (or variant) pointer argument.  

Some Hard Drive and File System benchmark tools

IOMeter

 

iozone

http://www.iozone.org

Document: http://www.iozone.org/docs/IOzone_msword_98.pdf
Manual: http://linux.die.net/man/1/iozone

Read, write, re-read, re-write, read backwards, read strided, fread, fwrite, random read/write, pread/pwrite variants

iozone -a | tee result.txt

iozone supports bunch of command line options. I summarized them in following table

Category Options Note
Auto mode -a: record size 4k - 16M, file size 64k - 512M.
-z: Used in conjunction with -a to test all possible record sizes. (Normally Iozone omits testing of small record sizes for very large files when used in full automatic mode. )
-A: more coverage
 
Test file -f filename: the name for temporary file under test.
-F fn1 fn2: # of files should be equal to # of processors/threads
 
Output -b filename: output of an Excel compatible file
-R: Generate Excel report.
 
Record size -r #: record size
-y #: minimum record size for auto mode
-q #: maximum record size for auto mode
 
File size -s #: size of the file to test
-g #: maximum file size for auto mode
-n #: minimum file size for auto mode
 
tests -i #: specifies which test to run
0=write/rewrite, 1=read/re-read, 2=random-read/write, 3=Read-backwards, 4=Re-write-record, 5=stride-read, 6=fwrite/re-fwrite, 7=fread/Re-fread, 8=mixed workload, 9=pwrite/Re-pwrite, 10=pread/Re-pread, 11=pwritev/Re-pwritev, 12=preadv/Re-preadv
One will always need to specify 0 so that any of the following tests will have a file to measure. This means -I 0 creates files used by following tests.
-i # -i # -i # is also supported so that one may select more than one test.
  -+p percent_reads: the percentage of threads/processes that will perform read testing in the mixed workload test case  
  -+B: sequential mixed workload testing  
throughput tests -t #: Run Iozone in a throughput mode.
-T: Use POSIX pthreads for throughput tests
This option allows the user to specify how many threads or processes to have active during the measurement.
processes/threads -l #: lower limit on number of processes to run
-u #: upper limit on number of processes to run
 
Timing -c: include close() in timing calculation
-e: include fflush(), fsync() in timing calculation
 
Other control -H #: Use POSIX async I/O with # async operations
-k #: Use POSIX async I/O (no bcopy) with # async operations.
-I: use direct I/O if possible
-m: use multiple buffers internally
-o: Writes are synchronously written to disk
-p: purges the processor cache before each file operation.
-W Lock file when reading or writing.
-K:  Inject some random accesses in the testing.
 
     

Examples

It's important to use -I option to turn on DIRECT I/O. Otherwise, linux's page caches (buffer cache) may give you ridiculous fast read speed.

  • Auto Mode

    • iozone -a -n 512m -g 1g

  • Single Test

    • Sequential write, Sequential reads
      iozone -r 64k -s 1g -b excel.xls -R -i 0 -i 1 -I
    • Use two processes to test sequential writes, sequential reads
      iozone -r 64k -s 1g -b excel.xls -R -i 0 -i 1 -I -l 2 -u 2
    • Use one and two processes to tests (two runs. In first run, one process is created. In second run, two processes are created)
      iozone -r 64k -s 1g -b excel.xls -R -i 0 -i 1 -I -l 1 -u 2
    • Random writes, Random reads
      iozone -r 64k -s 1g -b excel.xls-R -i 0 -i 1 -K -I
  • Throughput Test

    • iozone -t 2
      roughly equivalent to "iozone -l 2 -u 2"
  • Mixed Workload Test

    • iozone -r 64k -s 1g -b excel.xls -R -i 0 -i 8 -+p 50

Result visualization

./Generate_Graphs  result.txt

It will a directory for each operation (e.g. read, write, fread, fwrite). In each directory, there are two files generated - iozone_gen_out.gnuplot and <operation>.ps (this file is generated after you view the corresponding result using gnuplot).  Under the hood, it uses gnu3d.dem to render the data using Gnuplot after those data files are generated for all tested operations. You can call it directly without regenerating separate data files.

gnuplot gnu3d.dem

There are two more scripts that can be used for visualization - report.pl and iozone_visualization.pl.
When I tried to run them in Linux (using command ./report.pl result.txt), I had following error:

    -bash: ./iozone_visualizer.pl: /usr/bin/perl^M: bad interpreter: No such file or directory

You can use command perl report.pl result.txt to run it successfully.
The solution is to change those two files from dos type to unix type (mainly the new line character conversion).

sudo aptitude install tofrodos
fromdos report.pl
fromdos iozone_visualizer.pl

Now, you should be able to report.pl and iozone_visualization.pl directly. When they are run, a directory named 'report_result' is created. In the directory are Gnuplot scripts (*.do files) and PNG images for all tested operations. Those PNG images are generated by running those Gnuplot scripts. For example, read.png is generated by running "gnuplot read.do". A HTML page (index.html) is generated by iozone_visualization.pl which contains all of those PNG images in the same page.

Bonnie

http://www.textuality.com/bonnie/

Resources

Thursday, November 25, 2010

Permission of authorized_keys and private key file

~/.ssh/authorized_keys: should be 600

~/.ssh/id_dsa, ~/.ssh/id_rsa: must be 600

 

If permissions are not set correctly, the login process may fail without giving any useful information!

awk/gawk notes

       
RS Record Separator single character That character separates the records.
    regular expression Text in input matches reg exp separates records.
    null string Records are separated by blank lines. The  newline character always acts as a field separator, in addition to whatever value FS may have.
NR number of records seen so far    
FNR The input record number in the current input file    
ORS output record separator    
FS Field separator single character Fields are separated by that character
    single space fields are separated by runs of spaces  and/or  tabs  and/or  newlines.
    Null string each individual character becomes a separate field.
    regular expression  
OFS output field separator    
FIELDWIDTHS   a space separated list of numbers each field is expected to have fixed width. The value  of  FS  is ignored.  Assigning a new value to FS overrides the use of FIELDWIDTHS, and restores the default behavior.
NF number of fields   Decrementing NF causes the values of fields past the new value to be lost, and the value of $0 to be recomputed
IGNORECASE Controls the case-sensitivity of all regular expression  and  string  operations. non-zero
zero
non-zero: ignore case
       
Field Reference How to reference a field $1, $2, … $NF Access a field. Assigning  a value to an existing field causes the whole record to be rebuilt when $0 is referenced.
    $-1, $-2 fatal error
    non-existent fields For read, produce null-string. For write, 1)increase NF 2) create intervening fields with null string 3) $0 is recomputed
    $0 whole record.  assigning a value to $0 causes the record to be  resplit,  creating  new values for the fields.
CONVFMT    

A number is converted to a string by using the value  of  CONVFMT  as  a  format  string  for  sprintf(3), with the numeric value of the variable as the argument.  However, even though all numbers in AWK are floating-point, integral values are always converted as  integers.

OFMT      
All arrays in AWK are associative, i.e. indexed by string values.

i = "A"; j = "B"; k = "C"
x[i, j, k] = "hello, world\n"

key is "A\034B\034C" and value is "hello, world\n". Key test: val in array. for(val in array)…

Sunday, November 14, 2010

XLink

 http://www.xml.com/lpt/a/1038

Can be transformed to RDF?

One RDF use include to include other RDFs.

XLink and HLink: http://www.xml.com/lpt/a/1038

Enhancements to html link

  1. When to actuate the link
    In current html link impl, the link is actuated when it is clicked. In XLink and HLink, a link can be actuated when the page containing the link is loaded
  2. More effects when a link is actuated.
    embed, new, replace, etc.
  3. HLink supports creation of arbitrary link element.
    <hlink namespace="http://www.example.com/markup"
           element="home"
           locator="/"
           effect="replace"
           actuate="onRequest"/>
    <hlink namespace="http://www.example.com/markup"
           element="home"
           locator="/icons/home.png"
           effect="embed"
           actuate="onLoad"/>
    
    <home/>
  4. XLink supports creation of links among more than two resources.
  5. Add more metadata to links
  6. Links can be specified outside the linked resources.
    In HTML, users can only specify links within the source resource.
    When you write
     <a href=”destination.resource”>source</a>
    this piece of code must be located in the source html. In other words, the user cannot specify links among external resources.
    XLink adds this support.

 

METS, DIDL, ORE

http://www.oreillynet.com/xml/blog/2008/06/oaiore_compound_documents_draf.html

http://www.oreillynet.com/xml/blog/2008/05/bad_xml.html

http://www.dehora.net/journal/2008/06/18/dates-in-atom/

http://dret.net/netdret/docs/wilde-cacm2008-xml-fever.html

http://www.tbray.org/ongoing/When/200x/2006/01/09/On-XML-Language-Design

Google AppEngine mail test

Mail

How to test: http://aralbalkan.com/1311

Two bugs:

http://code.google.com/p/googleappengine/issues/detail?id=626
For this bug, you can upgrade your python to new version (2.5.4 and up).

http://code.google.com/p/googleappengine/issues/detail?id=1061

http://groups.google.com/group/app-engine-patch/browse_thread/thread/1662f95d9cacee24