#! contact | sed s@Fango@\ fanhoward@g | mail #!.com
#! fortune
I'm a proud Go contributor today. Wow.
May 10, 2012
走过这个城市
走在丝雨中,完全没有赶路的意思。一种安心,让我醉了。是陶醉在满足里。
走在从美术馆回来的路上,再次驻足在那家咖啡店的外面。已经要打烊了。女主人还在忙着。临街的座位,就是我带走满意的地方。
结帐时,我轻轻的问:“请问到底Forty Cafe里加了什么果实“?她答:”就只是这种高级的Espresso独特的香味。“。她朗朗的笑了一下。笑走了刚才满脸的倦意。是的,再用心做同一件事,久了也是无奈的疲态。但只需客人能略微的感动,就值回全心的付出了。
是一种什么样的感觉?那味道是直达舌底的,很熟悉又不能想起。一次次的抚起杯子,让那香气帮我回忆。慢慢一口一口的咽下,舌跟挥之不去的思绪。淡淡的甜,但带着浓郁。是了是了,是小时候留口水的烤地瓜的味道。
问好了去国美的路,目的是寻找这间咖啡店。刚好赶在饭口前,也因为没有预约,只能沿窗一排坐下。选好了生菜沙拉,清口。然后是一碗玉米粥配烤面包,适口。再下来主菜沙朗鹅肝。这时天已经完全黑了,店里也是满满的。但丝毫没有喧哗。大家都静静的交谈着,夹杂着轻碰玻璃杯的悦耳。忽然觉得腿有些麻了。原来是保持一个姿势太久了,太专注于沙朗的嫩和鹅肝的滑了。
”台中有些慢,还有些人情味在。“昨天回来时,计程车司机说。他以为我从台北过来,一路和我慢慢的聊着。讲着小时候还能搭公车,可现在捷运还不知何时建好,公车也越来越少。每台车都等在场里,等着附近的电话招呼。满街都是小摩托。那我那天从逢甲夜市出来,能在交通灯口拦下一辆空车,是幸运的了。
更幸运的是我能系上安全带。今天的计程车突然刹停。我的电脑包重重地甩到座位下。我也猛地从沉迷的书中抬起头来大喝怎么啦。司机大叔连声对不起,走下车,从路上拾起一个铁笼子,帮前面比卡的司机搬上去,却完全没有责怪的意思。这是个包容的城市,狭窄的路上忙乱的车流大家彼此迁就着。我也就不能再发作什么了。
但也不尽然,前天的路险可是让我Oh My God。同事开车送我,就惊见一辆红色小面包突然从拐角冲出来,闪过一个要横穿马路的摩托,擦着我们就跑远了。Oh My God。几乎就撞到那个万幸的骑士,几乎我就亲眼目击一桩惨祸。
今天的电视新闻,一辆花莲的旅游车,爬山路时换档不慎失去动力,载着13名南韩游客下滑翻覆到28米深的溪谷。万幸有树木拖住在8米处,又有一个管理员碰巧经过及时报警救援,得以保全每个人的生命。管理员的车载镜头录像了整个过程,使我们得以亲眼看到别人的不幸万幸,再庆幸。
又是一场不幸,雪山隧道猛烈燃烧一户人家化为碳与青烟,小孩无从辨认需DNA验证。为什么这两天台湾事故连连?其实这世界每时每刻都在重复着幸与不幸的故事,只是没有人强迫我们做我们旁观者。来了台湾才发现媒体上充斥的都是这些负面,或者是一些喧哗燥浮。不。这肯定不是全部。至少我亲眼亲身的是一个和善安逸的社会。
刚到台湾的第二天就去了日月潭。潭不大也不惊艳只是静静的绿水既不像日圆也没有月弯。和本岛游玩的人一样我租了一辆单车,就在和风煦日中远离了大团的游客。沿湖边一路骑行,轻松自在。偶尔停下看看,照个照片。偶尔猛骑一通,感受一下风的力度。不觉得3km的单车道到头要回头了,就不加思索的把车推上了山,准备来个33公里的环湖一圈。之前在游客中心问过要4个小时,也没注意到自己还穿着上班的牛仔T-shirt。等满身大汗气喘连连的爬到山顶,才后悔没换上短装。只好卷起裤腿坚持下去。还好下去的一路飞驰的痛快让我忘了刚刚的辛苦无奈。这样一路滑到12Km处,又是一段长长的上坡山路。37速的捷安特山地车也骑不动了,只能慢慢推行。沿路对面有些飞行而过的快乐骑士,还有个老太太冲我喊加油。感觉这到10Km处的两千米最为吃力。还好又是一路下坡直到玄奘寺渡口。当然想都没想做船横渡回去,还剩10公里已经想着老太太给的鼓励又上路了。这段上坡好辛苦,又要躲着大小汽车,又要担心峭壁的落石。还好带够了水一路不停地喝,又吃着飞机上留下来的杂果,熬过了午餐时段的饥饿,一路上上下下地到了尹达邵渡口。吃了山猪肉串和皮串,买了猫头鹰纪念,看看表已经骑了3个小时了,这剩下6公里的山路就坚持吧就坚持到了终点。车店的老板娘夸我厉害我只能苦笑着说很好玩然后等着晚上精疲力竭地和老婆Skype时再被夸赞损我go nuts。
-- fango
Aug 24, 2011
Almost renewed vim
I have been using Acme for many years and almost forgot how complex
the program editors could be. But last several months I have been
digging into headless embedded system, where only Vim is accessible
from serial line (ssh included). I used to be a man full of Vim, but
that days are long gone. Now I am a simple man who is contented with
simple things. So I cannot believe I have spent a whole week reading
'A Byte of Vim' while forcing myself using it for all my code works on
Windows/Linux/Macs, only because Plan9port stopped working after
upgrading to Mac OS X Lion. No. It is a mistake learnt the hard way.
Vim is gone away from me that no force can renew it back. Give up. I'm
a happy Acmer again, even that means I need to run Virtualbox ubuntu
on top of Lion, install Plan9port in order to get back my dear Acme.
Yes. Ubuntu is configured to log on directly into 'Recovery console',
where .profile runs rio & acme & google-chrome &, and xshove the
latter two to full screen, which I am used to alt-tab swapping them.
the program editors could be. But last several months I have been
digging into headless embedded system, where only Vim is accessible
from serial line (ssh included). I used to be a man full of Vim, but
that days are long gone. Now I am a simple man who is contented with
simple things. So I cannot believe I have spent a whole week reading
'A Byte of Vim' while forcing myself using it for all my code works on
Windows/Linux/Macs, only because Plan9port stopped working after
upgrading to Mac OS X Lion. No. It is a mistake learnt the hard way.
Vim is gone away from me that no force can renew it back. Give up. I'm
a happy Acmer again, even that means I need to run Virtualbox ubuntu
on top of Lion, install Plan9port in order to get back my dear Acme.
Yes. Ubuntu is configured to log on directly into 'Recovery console',
where .profile runs rio & acme & google-chrome &, and xshove the
latter two to full screen, which I am used to alt-tab swapping them.
Jun 9, 2011
SLEEF, an excursion on Go's math
A benchmark led to a contribution, and a discovery of sincos_amd64.s, further led to a pleasant sunny Sunday afternoon, reading the SLEEF paper. Thank you Mr. Naoki Shibata for your kind email.
How is that performed on ARM?
Misfortune struck. Just then I plugged out LAN cable from the pandaboard for Mac. When I plugged it back, it kept telling me files were read-only. Reboot without a proper shutdown (cannot as file system is readonly), rendered a corrupted EXT3-fs on my mmcblk0p2. All my math.Sqrt work were there, gone, luckily up to the cloud already.
DONT PANIC. Let's try e2fsck /dev/mmcblk0p2 first. Wow, after some clearing and fixing, SD is back to business.
The baseline, with Pow and Div.
sleef.BenchmarkSincos 1000000 2490 ns/op
sleef.BenchmarkMathSincos 5000000 442 ns/op
The optimized version, much faster, comparable to existing sin.go algorithm.
sleef.BenchmarkSincos 5000000 400 ns/op
sleef.BenchmarkMathSincos 5000000 443 ns/op
SLEEF has 128-bit SSE optimized code to obtain sin and cos at the same time, but sincos_amd64.s only uses 64-bit FPU, in effect it is just slightly optimized from SLEEF reference code in C.
I was not convinced the benefit of turning an optimized C to a handcraft assembly, merely to streamline the branches. Go should not be slow either. We just need numbers to say for itself.
A straightforward translation and benchmark shown 4x difference on my Mac Mini.
sleef.BenchmarkSincos 10000000 169 ns/op
sleef.BenchmarkMathSincos 50000000 39 ns/op
Turn math.Pow and div to mul, and inline the function, might speed it up. Let's try again:
sleef.BenchmarkSincos 50000000 51 ns/op
sleef.BenchmarkMathSincos 50000000 38 ns/op
The benefit is not so convincing to have an architecture dependent optimization. Compiler does a fairly good job already, and most importantly, it is trustworthy, and portable. Optimization without a benchmark always needs to be justified.
How is that performed on ARM?
Misfortune struck. Just then I plugged out LAN cable from the pandaboard for Mac. When I plugged it back, it kept telling me files were read-only. Reboot without a proper shutdown (cannot as file system is readonly), rendered a corrupted EXT3-fs on my mmcblk0p2. All my math.Sqrt work were there, gone, luckily up to the cloud already.
DONT PANIC. Let's try e2fsck /dev/mmcblk0p2 first. Wow, after some clearing and fixing, SD is back to business.
The baseline, with Pow and Div.
sleef.BenchmarkSincos 1000000 2490 ns/op
sleef.BenchmarkMathSincos 5000000 442 ns/op
The optimized version, much faster, comparable to existing sin.go algorithm.
sleef.BenchmarkSincos 5000000 400 ns/op
sleef.BenchmarkMathSincos 5000000 443 ns/op
Worth converting to SLEEF? It depends on another key factor - accuracy. That is the whole point of SLEEF - minimize the cancellation error to improve the accuracy, as demonstrated in the paper.
That accuracy holds for sincos, not exp. I had to reduce the tolerance from 1e-14 to 1e-6 to pass gotest. Exp.go is also slower than exp_386.s on my Linux 8g:
sleef.BenchmarkExp 20000000 117 ns/op
sleef.BenchmarkMathExp 50000000 69 ns/op
Ok. Last word to the world, the full set of benchmark for SLEEF sincos, asin/acos, exp and log, on Intel 32bit ubuntu , 64bit Mac OSX and ARM Cortex-A9.
Linux on Intel(R) Core(TM)2 Duo CPU T7300 @ 2.00GHz, bogomips : 3990.01
sleef.BenchmarkSincos 10000000 285 ns/op
sleef.BenchmarkMathSincos 50000000 55 ns/op
sleef.BenchmarkAsin 500000 3161 ns/op
sleef.BenchmarkMathAsin 20000000 103 ns/op
sleef.BenchmarkAcos 1000000 2017 ns/op
sleef.BenchmarkMathAcos 20000000 105 ns/op
sleef.BenchmarkExp 20000000 117 ns/op
sleef.BenchmarkMathExp 50000000 69 ns/op
sleef.BenchmarkLog 10000000 279 ns/op
sleef.BenchmarkMathLog 50000000 31 ns/op
Mac OSX on 2.4Ghz Intel Core 2 Duo
That accuracy holds for sincos, not exp. I had to reduce the tolerance from 1e-14 to 1e-6 to pass gotest. Exp.go is also slower than exp_386.s on my Linux 8g:
sleef.BenchmarkExp 20000000 117 ns/op
sleef.BenchmarkMathExp 50000000 69 ns/op
But faster than expGo on my ARM 5g:
sleef.BenchmarkExp 10000000 204 ns/op
sleef.BenchmarkMathExp 5000000 525 ns/op
It could be faster, as 5g does not do a good job at const folding yet, as demonstrated by NOT using T1 to T4 and calculating them in place instead:
sleef.BenchmarkSincos 1000000 1565 ns/op
sleef.BenchmarkMathSincos 5000000 442 ns/op
Enough benchmarking. My head is free-wheeling already. Go get a shot of Espresso and rest in pea's now!
Ok. Last word to the world, the full set of benchmark for SLEEF sincos, asin/acos, exp and log, on Intel 32bit ubuntu , 64bit Mac OSX and ARM Cortex-A9.
Linux on Intel(R) Core(TM)2 Duo CPU T7300 @ 2.00GHz, bogomips : 3990.01
sleef.BenchmarkSincos 10000000 285 ns/op
sleef.BenchmarkMathSincos 50000000 55 ns/op
sleef.BenchmarkAsin 500000 3161 ns/op
sleef.BenchmarkMathAsin 20000000 103 ns/op
sleef.BenchmarkAcos 1000000 2017 ns/op
sleef.BenchmarkMathAcos 20000000 105 ns/op
sleef.BenchmarkExp 20000000 117 ns/op
sleef.BenchmarkMathExp 50000000 69 ns/op
sleef.BenchmarkLog 10000000 279 ns/op
sleef.BenchmarkMathLog 50000000 31 ns/op
Linux on ARMv7 OMAP4 Pandaboard BogoMIPS : 2009.29
sleef.BenchmarkSincos 5000000 400 ns/op
sleef.BenchmarkMathSincos 5000000 443 ns/op
sleef.BenchmarkAsin 200000 12983 ns/op
sleef.BenchmarkMathAsin 1000000 1672 ns/op
sleef.BenchmarkAcos 200000 12890 ns/op
sleef.BenchmarkMathAcos 1000000 1705 ns/op
sleef.BenchmarkExp 10000000 204 ns/op
sleef.BenchmarkMathExp 5000000 531 ns/op
sleef.BenchmarkLog 5000000 662 ns/op
sleef.BenchmarkMathLog 5000000 418 ns/op
Mac OSX on 2.4Ghz Intel Core 2 Duo
sleef.BenchmarkSincos 5000000 51 ns/op
sleef.BenchmarkMathSincos 5000000 38 ns/op
sleef.BenchmarkAsin 200000 1122 ns/op
sleef.BenchmarkMathAsin 1000000 68 ns/op
sleef.BenchmarkAcos 200000 1107 ns/op
sleef.BenchmarkMathAcos 1000000 75 ns/op
sleef.BenchmarkExp 10000000 51 ns/op
sleef.BenchmarkMathExp 5000000 29 ns/op
sleef.BenchmarkLog 5000000 149 ns/op
sleef.BenchmarkMathLog 5000000 26 ns/op
Jun 2, 2011
Easy Going with ARM, square root up [3]
In my last post, we noticed that gc is still nearly twice as slow as the gcc for nbody benchmark. So, why? We put a simplest program under test.
Now compare the gcc's disassembly
Of course, SQRT in hardware already speeds up nbody benchmark by 7 times. ARM VFP also has ABS and NEG instruction. They are trivial at first glance, because what they do are simply clear, or negate the MSB (sign bit) of the operand register. Cannot this be done by simple AND or XOR? Yes. But we also need to know:
go@localhost:~/go/ex$ cat f.c
double mysqrt(double a){ return sqrt(a);}
main(){ mysqrt(0.1);}
go@localhost:~/go/ex$ cat f.go
package main
import "math"
func mysqrt(a float64) float64 {return math.Sqrt(a)}
func main(){ mysqrt(0.1)}
Now compare the gcc's disassembly
000083e0 <mysqrt>:
push {r7, lr}
sub sp, #8
add r7, sp, #0
strd r0, r1, [r7]
vldr d6, [r7]
vsqrt.f64 d7, d6
vcmp.f64 d7, d7
vmrs APSR_nzcv, fpscr
beq.n 8408 <mysqrt+0x28>
vmov r0, r1, d6
blx 8354 <_init+0x48>
vmov d7, r0, r1
vmov r2, r3, d7
mov r0, r2
mov r1, r3
add.w r7, r7, #8
mov sp, r7
pop {r7, pc}
00008418 <main>:
push {r7, lr}
add r7, sp, #0
add r1, pc, #16
ldrd r0, r1, [r1]
bl 83e0 <mysqrt>
mov r0, r3
pop {r7, pc}
to gc's:00010c00 <main.mysqrt>:
ldr r1, [sl]
cmp sp, r1
movcc r1, #180 ; 0xb4
movcc r2, #16
movcc r3, lr
blcc 14398 <runtime.morestack>
str lr, [sp, #-20]!
vldr d0, [sp, #24]
vstr d0, [sp, #4]
bl 280e4 <math.sqrt>
vldr d0, [sp, #12]
vstr d0, [sp, #32]
ldr pc, [sp], #20
00010c34 <main.main>:
ldr r1, [sl]
cmp sp, r1
movcc r1, #180 ; 0xb4
movcc r2, #0
movcc r3, lr
blcc 14398 <runtime.morestack>
str lr, [sp, #-20]!
ldr fp, [pc, #16]
vldr d0, [fp]
vstr d0, [sp, #4]
bl 10c00 <main.mysqrt>
ldr pc, [sp], #20
b 10c64 <main.main+0x30>
Russ is correct. Gcc uses intrinsics to generate VSQRT instruction in the code stream directly, while Gc uses a function call, which it seems expensive. Go will have function inlining, tough not available now. But look at A9 Neon MPE reference manual, VSQRTD uses 32 cycles (VDIV is 25, most others including VMUL are just 1 or 2 cycles), so the saving from function inlining may not help much.Of course, SQRT in hardware already speeds up nbody benchmark by 7 times. ARM VFP also has ABS and NEG instruction. They are trivial at first glance, because what they do are simply clear, or negate the MSB (sign bit) of the operand register. Cannot this be done by simple AND or XOR? Yes. But we also need to know:
- MPE has separate VFP and NEON (SIMD) block share the same 32 sixty-four bit register file.
- VFP has no logic operation.
- NEON has, but switching between VFP and NEON in A9, is expensive.
- Move back and forth between VFP and ARM core registers stalls both pipelines, should not be used in tight loop.
Reason No. 2 and 4 explains the reasons for VABS and VNEG, but not a strong one. I am happy to live without them.
Reason No. 3 is critical. FPU normally is mandatory for C and Go compilers (software emulation in both compiler and Linux kernel exists to help low end chips that without hardware floating point unit), but SIMD, is in the land of handcrafts, is used to speed up certain special operation, which in turn the special hardware logic tends to do much better job, thus renders SIMD a white elephant in the end. For example, NVidia is frequently challenged why not implement NEON in its cutting edge Tegra 2, the reason quoted usually is, with dedicate Audio/video codec and 2D/3D GPU, they don't see an immediate needs for NEON. Most other chips, including NVidia's next generation Tegra, have NEON, for programmers to waste time on.
Back to package math. I would like to port sincos_amd64.s to ARM VFP. That function can be used as the basis for other trigonometry functions. The algorithm, SLEEF, is intended for SIMD, but this implementation is not. It is a slightly optimized translation from C. So it would be straightforward to have a Go copy. But, as said in its comment, a VCMP would save a branch, always a win in modern deep pipelined CPU.
What does that mean? Dive into assembly again. A Go like this:
package main
func main(){
a, b := 1.0, 2.0
if a > b {
b = a
}
}
What does that mean? Dive into assembly again. A Go like this:
package main
func main(){
a, b := 1.0, 2.0
if a > b {
b = a
}
}
would 5g -S to this:
--- prog list "main" ---
0000 (f.go:3) TEXT main+0(SB),R0,$32-0
0001 (f.go:4) MOVD $(1.00000000000000000e+00),F4
0002 (f.go:4) MOVD $(2.00000000000000000e+00),F3
0003 (f.go:5) B ,5(APC)
0004 (f.go:5) B ,16(APC)
0005 (f.go:5) B ,7(APC)
0006 (f.go:5) B ,14(APC)
0007 (f.go:5) MOVD F4,F0
0008 (f.go:5) MOVD F4,F2
0009 (f.go:5) MOVD F3,F1
0010 (f.go:5) CMPD F3,F4,
0011 (f.go:5) BVS ,13(APC)
0012 (f.go:5) BGT ,6(APC)
0013 (f.go:5) B ,4(APC)
0014 (f.go:6) MOVD F2,F0
0015 (f.go:5) B ,16(APC)
0016 (f.go:8) RET ,
Lot of B for branches. Modern CPU puts lot of silicon to predict the branches, to squeeze out the pipeline bubbles caused by mis-prediction. But if we work on assembly, the code could be much optimized with conditional instructions. That could be a big deal for things like math.Sincos, because it is often used in tight loops for thousands of calculations.
Looking forward to another commit source-icide soon.
Looking forward to another commit source-icide soon.
May 31, 2011
Easy Going with ARM, Kung Fu Panda [2]
One of the fastest ARM board we can get today is $179 Pandaboard from Digikey. Do prepare to wait for months though.
I chose to install the headless image of ubuntu 11.04. Because my laptop has a SD slot on /dev/mmcblk0, the installation process was as smooth as my Teflon pan.
I chose to install the headless image of ubuntu 11.04. Because my laptop has a SD slot on /dev/mmcblk0, the installation process was as smooth as my Teflon pan.
- insert SD card to Laptop
- sudo umount /dev/mmcblk0
- sudo sh -c 'zcat ubuntu-11.04-preinstalled-headless-armel+omap4.img.gz > /dev/mmcblk0'
- sync
- insert SD card to pandaboard, plug power, LAN and USB-Serial cable in
- On the laptop terminal: TERM=vt100 minicom
- turn on the pandaboard, on minicom after the uboot comes the standard ubuntu installation process, and finally gives us a shell prompt.
The Go installation is the same as in my last post, only much faster. Now comes to the benchmark, how does 1GHz dual core ARM A9 compare to my 2GHz dual core Intel x86, and 1GHz single core ARM A8?
Not surprisingly, the number of cores does not count here since no parallel processing is benchmarked. For string processing 1GHz A9 is slightly faster than A8 , but still more than twice slower than 2GHz x86 core. A9's VFP has been greatly improved, 5x faster than A8 by gcc.
But surprisingly, and I am astonished to see, for the floating point crunching on A9, gc is 11x slower than optimized gcc. This is very unfortunate because what I am interested in Go on ARM is OpenGL ES, which is all about matrix operations on floating points.
[update] 11x slower is caused by unoptimized pkg/math/sqrt.go, since ARM VFP has VSQRT instruction, it should not be hard to speed it up.
[update 2] I made it. Now it is 7x faster
nbody -n 50000000
gcc -O2 -lm nbody.c 71.40u 0.00s 71.43r
gc nbody 120.93u 0.00s 120.94r
gc_B nbody 119.78u 0.00s 119.80r
[update] 11x slower is caused by unoptimized pkg/math/sqrt.go, since ARM VFP has VSQRT instruction, it should not be hard to speed it up.
[update 2] I made it. Now it is 7x faster
nbody -n 50000000
gcc -O2 -lm nbody.c 71.40u 0.00s 71.43r
gc nbody 120.93u 0.00s 120.94r
gc_B nbody 119.78u 0.00s 119.80r
[/update2]
go@localhost:~/go/test/bench$ cat /proc/cpuinfo
Processor : ARMv7 Processor rev 2 (v7l)
processor : 0
BogoMIPS : 2009.29
Processor : ARMv7 Processor rev 2 (v7l)
processor : 0
BogoMIPS : 2009.29
processor : 1
BogoMIPS : 1963.08
Features : swp half thumb fastmult vfp edsp thumbee neon vfpv3
Hardware : OMAP4 Panda board
BogoMIPS : 1963.08
Features : swp half thumb fastmult vfp edsp thumbee neon vfpv3
Hardware : OMAP4 Panda board
go@localhost:~/go/test/bench$ gomake timing
./timing.sh
reverse-complement < output-of-fasta-25000000
gcc -O2 reverse-complement.c 7.88u 1.55s 9.54r
./timing.sh
reverse-complement < output-of-fasta-25000000
gcc -O2 reverse-complement.c 7.88u 1.55s 9.54r
gc reverse-complement 18.36u 1.91s 20.29r
gc_B reverse-complement 17.75u 2.08s 19.85r
gc_B reverse-complement 17.75u 2.08s 19.85r
nbody -n 50000000
gcc -O2 -lm nbody.c 71.40u 0.00s 71.41r
gc nbody 862.53u 0.02s 862.78r
gc_B nbody 865.00u 0.05s 865.28r
gcc -O2 -lm nbody.c 71.40u 0.00s 71.41r
gc nbody 862.53u 0.02s 862.78r
gc_B nbody 865.00u 0.05s 865.28r
May 30, 2011
Easy Going with ARM
How easy would it be to Go with ARM? It depends, from super difficult to duper easy. This is the latter case. Just got one from Mouser for US$149 with free FedEx to Singapore. Opened the box, plugged the cables in, inserted the SD card, and pressed the power switch.
As advertised, I should expect an instant boot-up. But nothing happened. Dead on arrival? Scratched my head. Where's the manual? Searched the box again. Not found. Looked at the board and thought, should I insert the SD or microSD? No harm to try.
Sliced the microSD from the SD sleeve, put in that slot, powered up. Bingo. AND as I idly lying back on my armchair and glanced the box again... There you are. The manual and CD are just there, pasted on the back of the box cover!
From then on, it could not be easier to just follow the normal procedure to have a Go on the pre-installed Ubuntu Lucid on the IMX53 Starter-Kit.
- sudo apt-get update
- sudo apt-get install mercurial bison ed gawk gcc libc6-dev make
- hg clone -u release https://go.googlecode.com/hg/ go
- cd go/src; ./make.bash
It just needs much longer time (an hour?), because it only runs on a 1GHz ARM Cortex-A8, with 1GB DDR3 RAM, and worst of all, SD is way too slow. Great news is this pixie has SATA port. Next time will try find a spare hard disk and see how speedy it could go.
How slow is it? Here's the go/test/bench on my PC and ARM. I just picked two to show the string processing and floating point crunching, and cpuinfo is shorten to show the relevance only.
#! cat /proc/cpuinfo
model name : Intel(R) Core(TM)2 Duo CPU T7300 @ 2.00GHz
stepping : 11
cpu MHz : 800.000
cache size : 4096 KB
bogomips : 3990.32
#! make timing
reverse-complement < output-of-fasta-25000000
gcc -O2 reverse-complement.c 1.43u 0.23s 1.67r
gc reverse-complement 2.72u 0.26s 3.00r
gc_B reverse-complement 2.64u 0.31s 2.95r
nbody -n 50000000
gcc -O2 -lm nbody.c 27.78u 0.00s 27.92r
gc nbody 54.01u 0.00s 54.06r
gc_B nbody 52.18u 0.00s 52.27r
lucid@lucid-desktop:~$ cat /proc/cpuinfo
Processor : ARMv7 Processor rev 5 (v7l)
BogoMIPS : 999.42
Features : swp half thumb fastmult vfp edsp neon vfpv3
CPU implementer : 0x41
CPU architecture: 7
Hardware : Freescale MX53 LOCO Board
lucid@lucid-desktop:~/go/test/bench$ gomake timing
reverse-complement < output-of-fasta-25000000
gcc -O2 reverse-complement.c 8.00u 1.14s 10.27r
gc reverse-complement 22.97u 1.26s 27.62r
gc_B reverse-complement 22.09u 1.52s 30.77r
nbody -n 50000000
gcc -O2 -lm nbody.c 316.08u 0.40s 389.46r
gc nbody 645.32u 686.06s 1843.30r
gc_B nbody 653.49u 640.93s 1373.40r
Not surprisingly, Cortex-A8 VFP is known to be very slow, due to its non-pipeline architecture. But I don't know it is so tortoise-like. Will be Cortex-A9 much better? Once I get my OMAP4430 panda board work, I will report it here.
Just to confirm I was using VFP not the soft-float, I created each version and compare:
$ 5l -F -o nbody.arm5 nbody.5
$ 5l -o nbody.arm6 nbody.5
$ time ./nbody.arm6 -n 50000
real 0m1.316s
user 0m1.310s
sys 0m0.000s
$ time ./nbody.arm5 -n 50000
real 0m30.788s
user 0m29.830s
sys 0m0.000s
By the way, the proper way to run glib is via pkg-config as below. But I gave up trying to fix run().
run 'gcc -O2 `pkg-config --cflags glib-2.0` k-nucleotide.c `pkg-config --libs glib-2.0`' a.out <x
(the mouse is for illustration only, not come with the board)
May 18, 2011
Black Perl
偶然看到一篇 Perl 语言的诗篇, 据说可以在 Perl 3 上编译. 作者 Larry Wall, 语言学家, Perl 的发明者和导师. 神奇的是我熟悉的 Perl 很像 C, 变量全都有钱, 而这个诗篇, 除了 die, 没一点 Perl 的特征. 所以人称 Perl "There's more than one way to do it", 果不其然.
BEFOREHAND: close door, each window & exit; wait until time.
open spellbook, study, read (scan, select, tell us);
write it, print the hex while each watches,
reverse its length, write again;
kill spiders, pop them, chop, split, kill them.
unlink arms, shift, wait & listen (listening, wait),
sort the flock (then, warn the "goats" & kill the "sheep");
kill them, dump qualms, shift moralities,
values aside, each one;
die sheep! die to reverse the system
you accept (reject, respect);
next step,
kill the next sacrifice, each sacrifice,
wait, redo ritual until "all the spirits are pleased";
do it ("as they say").
do it(*everyone***must***participate***in***forbidden**s*e*x*).
return last victim; package body;
exit crypt (time, times & "half a time") & close it,
select (quickly) & warn your next victim;
AFTERWORDS: tell nobody.
wait, wait until time;
wait until next year, next decade;
sleep, sleep, die yourself,
die at last
# Larry Wall
再找找看, 早有人作了 C 的诗仙. 可编译, 可运行, 更是越读越有趣!
char*lie;
double time, me= !0XFACE,
not; int rested, get, out;
main(ly, die) char ly, **die ;{
signed char lotte,
dear; (char)lotte--;
for(get= !me;; not){
1s - out & out ;lie;{
char lotte, my= dear,
**let= !!me *!not+ ++die;
(char*)(lie=
"The gloves are OFF this time, I detest you, snot\n\0sed GEEK!");
do {not= *lie++ & 0xF00L* !me;
#define love (char*)lie -
love 1s *!(not= atoi(let
[get -me?
(char)lotte-
(char)lotte: my- *love -
'I' - *love - 'U' -
'I' - (long) - 4 - 'U' ])- !!
(time =out= 'a'));} while( my - dear
&& 'I'-1l -get- 'a'); break;}}
(char)*lie++;
(char)*lie++, (char)*lie++; hell:0, (char)*lie;
get *out* (short)ly -0-'R'- get- 'a'^rested;
do {auto*eroticism,
that; puts(*( out
- 'c'
-('P'-'S') +die+ -2 ));}while(!"you're at it");
for (*((char*)&lotte)^=
(char)lotte; (love ly) [(char)++lotte+
!!0xBABE];){ if ('I' -lie[ 2 +(char)lotte]){ 'I'-1l ***die; }
else{ if ('I' * get *out* ('I'-1l **die[ 2 ])) *((char*)&lotte) -=
'4' - ('I'-1l); not; for(get=!
get; !out; (char)*lie & 0xD0- !not) return!!
(char)lotte;}
(char)lotte;
do{ not* putchar(lie [out
*!not* !!me +(char)lotte]);
not; for(;!'a';);}while(
love (char*)lie);{
register this; switch( (char)lie
[(char)lotte] -1s *!out) {
char*les, get= 0xFF, my; case' ':
*((char*)&lotte) += 15; !not +(char)*lie*'s';
this +1s+ not; default: 0xF +(char*)lie;}}}
get - !out;
if (not--)
goto hell;
exit( (char)lotte);}
BEFOREHAND: close door, each window & exit; wait until time.
open spellbook, study, read (scan, select, tell us);
write it, print the hex while each watches,
reverse its length, write again;
kill spiders, pop them, chop, split, kill them.
unlink arms, shift, wait & listen (listening, wait),
sort the flock (then, warn the "goats" & kill the "sheep");
kill them, dump qualms, shift moralities,
values aside, each one;
die sheep! die to reverse the system
you accept (reject, respect);
next step,
kill the next sacrifice, each sacrifice,
wait, redo ritual until "all the spirits are pleased";
do it ("as they say").
do it(*everyone***must***participate***in***forbidden**s*e*x*).
return last victim; package body;
exit crypt (time, times & "half a time") & close it,
select (quickly) & warn your next victim;
AFTERWORDS: tell nobody.
wait, wait until time;
wait until next year, next decade;
sleep, sleep, die yourself,
die at last
# Larry Wall
再找找看, 早有人作了 C 的诗仙. 可编译, 可运行, 更是越读越有趣!
char*lie;
double time, me= !0XFACE,
not; int rested, get, out;
main(ly, die) char ly, **die ;{
signed char lotte,
dear; (char)lotte--;
for(get= !me;; not){
1s - out & out ;lie;{
char lotte, my= dear,
**let= !!me *!not+ ++die;
(char*)(lie=
"The gloves are OFF this time, I detest you, snot\n\0sed GEEK!");
do {not= *lie++ & 0xF00L* !me;
#define love (char*)lie -
love 1s *!(not= atoi(let
[get -me?
(char)lotte-
(char)lotte: my- *love -
'I' - *love - 'U' -
'I' - (long) - 4 - 'U' ])- !!
(time =out= 'a'));} while( my - dear
&& 'I'-1l -get- 'a'); break;}}
(char)*lie++;
(char)*lie++, (char)*lie++; hell:0, (char)*lie;
get *out* (short)ly -0-'R'- get- 'a'^rested;
do {auto*eroticism,
that; puts(*( out
- 'c'
-('P'-'S') +die+ -2 ));}while(!"you're at it");
for (*((char*)&lotte)^=
(char)lotte; (love ly) [(char)++lotte+
!!0xBABE];){ if ('I' -lie[ 2 +(char)lotte]){ 'I'-1l ***die; }
else{ if ('I' * get *out* ('I'-1l **die[ 2 ])) *((char*)&lotte) -=
'4' - ('I'-1l); not; for(get=!
get; !out; (char)*lie & 0xD0- !not) return!!
(char)lotte;}
(char)lotte;
do{ not* putchar(lie [out
*!not* !!me +(char)lotte]);
not; for(;!'a';);}while(
love (char*)lie);{
register this; switch( (char)lie
[(char)lotte] -1s *!out) {
char*les, get= 0xFF, my; case' ':
*((char*)&lotte) += 15; !not +(char)*lie*'s';
this +1s+ not; default: 0xF +(char*)lie;}}}
get - !out;
if (not--)
goto hell;
exit( (char)lotte);}
Aug 4, 2005
Form is liberating

昨天又重读<人月神话>, 这次是英文版. 概念都已经很熟悉, 但原汁原味的感觉仍带来很多感悟. "Form is liberating", 如何理解? 为什么古诗词都要有固定格式? 作曲要讲韵律? 建筑要有风格? 而这些又能让大师们尽情创作传世精品? 中文版的翻译为"没有规矩, 不成方圆". 翻译不错, 但那是常用来教训孩子的话, 听起来刺耳, 不如就直译 "格式带来解放". 另 'Sharp tools' 译为'干将莫邪', 不知是指人指物, 不如直译为 '精工利器'.
话归正题. 和建筑业一样, 设计师的灵魂体现在整体的外观, 也渗透在每一个结构. 他必须了解每一种材料对整体效果和结构的影响, 尽管不需亲自施工, 但他动手的作品应能成为其他工匠的参照. 设计师定义好了风格, 其他工匠和设计师才可以解放思想, 创造又不破坏作品的完整.
和建筑业不同, 计算机系统不需要工人. 每个人都参与创造, 所以每人都要保证自己的作品符合系统的风格. Inferno 和 UNIX, Plan9 一脉相承, 大师 Dennis , Ken, Rob 等的风格就是我们要遵循的规范, 他们的源代码就是我们的典范, 我们的创造力不可破坏 Plan9 三原则, 即:
1. 所有资源都是文件.
2. 文件使用统一协议.
3. 分立资源统一在私有命名空间.
Aug 3, 2005
Small is beautiful

如果把软件比作衣服, 那我们还生活在19世纪. 为了向后兼容, 或传统或规范, 我们必须穿上三件套, 戴上高筒帽, 拄上文明棍儿. 更糟, 软件衣服在保持向后兼容的同时, 还会照顾人们的特殊需要, 有人爱远游, 有人爱打猎, 于是三件套有了内置背包和外挂枪套. 因为要大量生产来降低设计成本, 于是没有了正好合适的衣服, 于是每个人被迫花高价穿着这些去上班. 更糟, 常有人忘记了钱包是随手放进了哪个背包或枪套再也找不到, 而厂商已经声明 AS IS 概不负责, 还有人写傻瓜书 1234 教人们如何定位口袋, 破解扣子, 安放钱包, 再如何查寻是放在哪个口袋...
还好人们解放了自己的衣服, 在家老头衫, 上班牛仔裤, 海边比基尼. 大厂小厂, 大店小店, 衣服千变万化, 供人们随意搭配. 人们喜欢的是合适的衣服, 能自由搭配, 价格也不能贵.
小而专, 是 UNIX 传统的设计理念, 但当更多的人开始编写 UNIX 程序时, 从其它系统得来的思维习惯开始污染她, 于是有了网络系统和X 窗口系统和更多庞大复杂试图自成系统的系统, 每个系统不但内部结构复杂, 对外的接口也各自为政. 系统间不再能轻易的互联, 系统不再能轻易的裁减, 不再有刚好合适的系统, 而只是一大堆不同系统奇怪的堆积.
Plan9 重新回归了小而专的理念. 文件和管道是程序互联的基础. Plan9 进一步把所有可用的资源文件化, 又用频道使多个进程轻易互联. 这些使Plan9 程序开发和维护变得简单.
Inferno 进一步将程序代码变为可随时共享的资源. 专一的格式免去为不同体系结构重新编译的麻烦, 也保证了新的体系结构能立即使用现有代码.
Inferno 模块化的结构使每一模块只专注做好一件事, 用户又能自由组合模块, 裁减系统使之刚好合适. 统一的外部文件接口, 使用户能轻易的组织系统. 甚至只使用Shell 也可搭建完整的网络化图形界面程序!
Inferno 是一套很小很纯的系统, 不要污染她, 用心体会她的设计理念: 小即是美!
Aug 1, 2005
The revolution of evolution

从进化的角度看, Blackfin 已不再是一颗普通的DSP 或MCU. 它已经将这两个品种的优点揉合在一起. 定名为 MSA, 即综合信号体系, 具有非常强大的数字运算能力, 又有灵活的内存和中断服务. 统一精简的指令集使 Blackfin 无论从性价比还是使用方便上, 都超过分立的DSP和MCU.
为 了达到最大的定点运算性能, Blackfin 使用了十级流水线, 单周期的乘法指令, 双乘加器, 硬件循环和宽达160位的片内数据通路, 可以同时读取四条16位指令, 四个16位操作数, 将其两两相乘并累加, 再将一个32位的结果写回, 比如下面 ADI VDSP++ 的汇编指令:
a1+=r0.l * r2.h, a0+=r0.h * r2.l || r2.l = W[i2++] || [i1--] = r3
或对应的 Inferno 的汇编指令:
a2 + =r0 * r2; r2.l.w = (i2)+; (i1)- = r3;
这种体系结构为大量的DSP运算提供了必要的支持, 这不是在现有的MCU 如ARM上增加几条优化指令就能做到的.
同时, Blackfin 具备12通道的DMA, 两层15级中断, MMU提供的内存保护和缓存, 这些为编写复杂的控制软件和操作系统提供了可能. 已有大量的实时操作系统直接运行在 Blackfin 上, 包括Linux 2.6, 将 Inferno 移植到 Blackfin 上也没有任何技术问题.
其集成在片上的SDRAM控制器, 高速串口, 并口, 时钟, 甚至以太网控制器, 也为连接外设搭建完整系统提供了极大方便.
更让人钦佩的是, Blackfin 出色性能的背后, 是低廉的成本和很低的功耗. 这和现时片面追求高性能成为鲜明对比. 我们需要的是能满足基本性能的最便宜最省电的芯片, Blackfin 是最佳选择.
Jul 31, 2005
Beauty is in the eyes of beholders

如果用一句话总结 Inferno 的特点, 可以是: 一套精练完备的实时多任务操作系统, 及其网络文件接口, 图形界面, 虚拟机, 仿真器和交互式交叉编程环境.
而 Blackfin 的特点可总结为: 一个由精简高效的数字信号处理结构, 及其丰富的外设接口所构成的低价芯片.
Blackfin 没有使用复杂的MMU, 可以有效的降低成本, 代价是失去内存保护和切换, 使其不能安全的运行大量现有程序, 只能作为专用芯片使用. 如果存在一种操作系统, 不需 MMU 也可安全加载运行用户程序, 便可充分利用其极高的性价比, 构成新一代的通用计算机, 性能不减, 而成本远低于PC.
Inferno 就是这种操作系统. 除了内核和驱动程序是用 C 实现, 所有的应用程序用 Limbo 编写, 并编译成中间码, 由 Dis 虚拟机执行. Dis 提供了内存保护和对应机器码的翻译, 使 Dis 代码能直接安全运行在所有 Inferno 支持的系统上, 这包括 PC, Mac 和 UNIX, Plan9 上的仿真器, 和运行在ARM, x386, PowerPC, MIPS 上的 Inferno 操作系统.
Limbo 是一种综合了 C 和 Pascal 优点的模块化多线程编程语言. 没有指针, 编译和运行态的类型检查, 垃圾回收 (Gabage Collection), 提供了安全有效的内存管理. 模块(module)实现和加载提供了结构化的程序管理和运行环境的动态改变. 通讯顺序进程(CSP)使用频道(channel) 的发送和接收简单实现了多线程的管理.
在 Inferno 系统中, 所有的可用资源都用文件代表, 包括内核的, 外存的, 和网络上的数据, 状态和控制. 完成文件操作的协议是 Styx, 定义了文件的创建, 打开和关闭, 游标的复制和移动, 以及数据和属性的读写. 用户和程序可以将用到的文件名组织成自己私有的目录, 称为命名空间(namespace).
学习 Inferno, 只需从 www.vitanuova.com 下载安装. 阅读理解其中的代码和文档, 再自己编写一些程序来掌握和精通.
学习 Blackfin, 可以从 www.digikey.com 购买一套低价的 Blackfin STAMP 实验板. 后文将提到如何在 PC 上编译 Inferno Blackfin, 并安装到这个 STAMP 实验板上.
暂时没有 Blackfin 实验板也没关系, 在 PC 上可以作大部分的仿真操作.
Subscribe to:
Posts (Atom)


