百萬企鵝 A Million Penguins
百萬企鵝是一個嘗試由網路使用者共寫的小說撰寫計畫, 概念顯然地是基於 Wikipedia, 當初看到這個新聞 (Knews) 時, 除了覺得很有趣之外, 其實還對新聞中的一段敘述感到有意思與不解 :
企鵝出版集團強調,這項計畫並非在發掘新作家,所以每位「作者」以 250 字為限,他們呼籲網友要有雅量,尊重別人的國籍、文化及性別,因為每個人的創作都可能會被改動,受不了自己的文章被隨意刪改的人也許不參加為妙。(引用自 Knews )
這是被討論很久的問題, 特別是當初 Wikipedia 計畫開始推廣後, 相對的問題就不曾停止過, 例如 中文 Wikipedia 上針對中華民國的 頁面, 因為爭議太多, 持相異意見的會彼此胡亂刪減修改對方內容, 一直以來都是處於 Semi-Protection 模式下, 幾乎沒辦法對內容做修改, 只能等待討論區有共識之後再說, 換而言之, 在 Wikipedia 上針對此問題的解決是利用 Locking. 真跟上面百萬企鵝新聞稿中的敘述其實很像.
然而, 百萬企鵝何苦要遵循此種模式 ?
回到百萬企鵝計畫的目的, 在他們的 About 中, 首先出現的是粗體字的
Crowdsourcing. The Wisdom of the crowds. Social networking. Collaborative enterprise.
可見百萬企鵝與 Wikipedia 的共同特點, 但是他們決定性的不同卻在於企望達到的目標上. 再參考 About 內的另一段話 :
But what about the novel? Can a collective create a believable fictional voice? How does a plot find any sort of coherent trajectory when different people have a different idea about how a story should end – or even begin? And, perhaps most importantly, can writers really leave their egos at the door?
小 說往往並非事實, 它是 fictional voice 的, 不若 Wikipedia 是以共同建構知識庫為目標. 在 Wikipedia 中, 因為有事實存在, 因此難以容許參雜個人觀感推斷的敘述在內, 然而許多事實卻無法被證明, 只能靠著佐證的史料進行推斷, 由此矛盾就生. 這是短時間內不可能解的問題, 除非我們再度回到過去的獨裁帝制. 百萬企鵝卻不是這樣, 它的計畫要產出的成果是小說, 小說是不是跟事實相關其實根本不重要, 甚至許多地方的邏輯不通其實也不甚重要, 如果我們不期待名家大作, 那麼能有條有理, 有頭有尾地說完一個故事, 其實才是最重要的.
同時, 我覺得這段話裡重新說明了一個重要的事實 : 不同的人本來就會有不同的想法, 只要說的通, 要怎麼想是每個人的自由. 或許小說作者有他想要表達的意念在他的作品中, 但是究竟閱讀者會怎樣解讀, 那其實不是作者能夠控制的事情, 某個角度來說, 作者也只是一個讀者而已. 過去受限於人類所能做到的能力, 小說都是關起門來寫的. 百萬企鵝如果成功, 象徵的其實是打破這個藩籬, 並可以進一步尋求有效 ( efficient ) 的 process 出現. 類似的現象已經在 software domain 出現過了, Eric S. Raymond 著名的 The Cathedral and the Bazaar 十分清楚地說明了此一現象的轉移, 而現在, OSS community 已經在尋找與嘗試 efficient development process 了 ( 但這將會是條非常漫長的路 ) .
百萬企鵝其實應該回到最基本想證明的問題上, 在這樣的環境下, 幾百萬人是不是能夠合作共寫出一部小說其實並不是一個真正需要被解決的問題. 幾百萬人如何能夠合作共同寫出幾萬部小說才是真正應該要被解決的問題. 而百萬企鵝最後產出的那部小說, 其實就是這幾萬部小說中, 最受所有人喜愛的那一本.
雖 然我還不知道這個問題要怎樣解( 要是知道就去寫 paper 開公司了 XD ), 但是可以想見的或許把 content 跟 logic 分開會是必要的, 同時不同版本的 content, logic 是可以同時存在的. 以下圖來說明, 一部小說的內容可以分成數段的 content piece, 靠著 logic 將其連結起來. 如此一來, 每個人可以貢獻自己的想法, 撰寫其中一個 content piece 就好, 不必寫出一整部, 或是較大的區段, 同時結合自己的, 別人的 content pieces, 可以用自己特定的 logic 將其串起來, 但不需要串起全部的 content pieces, 只需要串起說的通的就可以. 從而網路貢獻者又可以分為 context piece contributor 以及 logic contributor 兩類.
同時 content piece 在一定的授權下是可以進一步被拿去修改擴充的, 因此不必有所謂修改刪除別人的文章之情形出現, 而是多重版本可以同時存在, 例如下圖中的 context piece B'' 就是由 content piece B' 修改來, 而 content piece B' 又是由 content piece B 而來. 另外 logic 本身也可以是被重新利用修改的, 例如 Logic One' 就是由 Logic One 修改而來, 替換了最前面的兩個 context piece 連結.
而 logic 與 context pieces 間的關係, 藉由 logic 與 content piece 的分離, 他們之間的關係是 flexible 的, 這樣的想法主要來自於 In-Code Modeling .
如 此一來就不會有百萬企鵝或是 Wikipedia 中, 因為彼此不滿意內容而刪減, 造成爭執的情況出現, 而且各種可能的串聯邏輯可以出現, context pieces 以及 logic 可以很方便地被 reuse / re-use, 整體的 productivity 因而可以提升, 多線劇情小說 (例如 MLNR, 此計劃似乎停擺很久了 ) 的想法也將可以被實現, 這基本上也是 component-based software development (CBSD) 的理想. 當然, 適當的配套措施是要有的, 例如怎樣比對相似的內容, 減少 context pieces 的數量, 避免因為大量的 content pieces 造成 logic contributor 反而需要花費大量時間閱讀可用的 content piece material. 或是適當的 search mechanism 可以被支援, 自動依照 logic contributor 提供的一些線索, 找出適用的 content pieces. 這其實都相似於 CBSE 對於 component 的處理.
或許我之前那邊 Paper Writing Process 也可以朝此方向去解決.
下午1:00 | 標籤: idea | 0 Comments
Attacking wretch.cc with DDOS
之前看到 runtime@ptt.cc 發布的連連看服務 ( Wretch Relation Map, 不確定此連結會活多久), 其實跟我之前的 Author Net 是一樣的東西. 其實之前也有想過對無名做這種事情, 只是對我實在沒什用處, 加上還要提供 Server 實在太麻煩了, 如果是我應該會寫 client 端程式, 然後發佈出去. 但是這個要考量的就多了, 例如會不會因此變相形成對於無名的 DDOS 攻擊, 到時候帳算到我頭上就不好了.
不過這樣的 DDOS 攻擊蠻弱的就是了, 只要無名改一下 friends link 的呈現馬上就無效了. 但是他也很難就改成無法直接取得的方式, 否則要不就是 usability 降低, 要不就是 perfomance 降低.
真 好奇像是這種 popular site, 怎樣面對合理服務要求的大量 requests, 即便 request 量幾乎等同於 DDOS attack. 當然利用各種 queue algorithm 是一種解法, 但是對於付費者來說, 要跟一般使用者等待相同的時間嗎 ? 提供 service 的 site 怎樣區別一般使用者跟付費使用者 ( 在可以進行 login 動作之前 ), 有可能在 socket connection 建立時就辨認嗎 ? 利用 cookie 是否可靠 ? 但是 cookie 的使用僅在使用者經常使用固定電腦實際較為有效而已.
什麼樣的 services 應該是 open (例如 friends list 察看), 什麼樣的應該是 protected, 這之間跟 site usability 的 trade-off 又是如何, 真是個有趣的問題.
中午12:55 | 標籤: idea, security | 0 Comments
Paper Writing Process
之前在趕 TSE special issue on SESS 的 paper, 跟兩位學長合作寫, 感覺寫的好累.
一方面是時間很趕, 一方面是不久前才完成 ESEM 2007 的 conference paper, 緊接著就是 TSE, 感覺有點疲累, 加上 Lab. 內部的其他小計畫, 還有我的 group meeting presentation 題目未定, 事情好多阿.
寫 Paper 真的需要這麼累嗎 ? 怎樣可以較容易達成 paper writing 的分工 ?
照 現在的運作, 大概是在 paper 的架構訂出來之後就分下去寫了, 對比到 software development 中大概是 architecture 大致確立的階段就分了, 換句話說在 design phase 事實上還有相當大的 variation 存在, 每個人的想法跟理解未必相同, 寫出來的東西在做 integration testing 時需要花的功夫相當大.
寫 paper 究竟能不能像寫 software 一樣, 可以有一個較為嚴謹的 process ?
針對這個問題, 有幾個其實較為明顯可以去思考的點 :
- Clear responsibility. 在 規劃 software architecture 時, 利用 CRC card 之類的去確認每個 object (under OO paradigm) 的 responsibility 是常見的作法, 對於 architecture 中的各部份之 responsibility 都必須確定, 才知道後續的 design 需要滿足什麼需求, 也不需要多做不必要的事情. 因此是否 paper 的各段落也可以用類似的方法, 確定各段落的 responsibility, 甚至是 non-functional 的條件
- Regular review. 有 點像 iterative 的開發方式, 在分工撰寫時, 利用 regular 的 meeting, 彼此 reivew 寫好的部份, 確認彼此所寫的仍然忠於原先的 responsibility 規劃, 同時 writing style 可以適度做修改. 關於 writing style, 畢竟 natural language 不像 programming language, 可以制定非常準確的 writing style, 但是以 academic writing 的 context 來說, 應該還是可以訂出一些 style, 這可能要參考一些 academic writing 的書
- Version control. 如果把 paper 視為 software, 用 version control system 去管理 changes 應該是沒有問題的事情, 只是過去習慣用一個 latex file 紀錄全部文字的作法就要改, 但是基本上這不是什麼大問題. 借由 version control system, 所有的 changes 都更容易進行, 也未必必須要負責該部分的人才能修改. 老師也比較容易看到所有的修改紀錄. 唯一的缺點可能是 security 的問題, 最好該 version control system 架設在實驗室內部網域應該就可以了.
中午12:51 | 標籤: idea, research | 0 Comments
Component Modeling for Classification and Retrieval : A Comparison of Previous Works
Recently, I am working on a small Lab. project about a novel approach for component storage and retrieval to assist component reuse in CBSD. In the domain analysis phase, some component modeling approaches were surveyed and compared. They were organized (ordered by publication year) in following table for anyone interested.
| Works | Modeling Approach | Modeling Target | Classification Base | Retrieval Base |
| [Ostertag1992] | Graph-based | Functional Property / Component Feature | AI-based / Component Feature ( Feature Distance ) | |
| [Frakes1994] ( Survey Paper ) | Keyword Indexing | Component Characteristics | Controlled Vocabulary / Uncontrolled Vocabulary | Keyword/ Hypertext |
| [Zaremski1995] | Signature-based | Function Types / User-defined Types | Signature Matching | Signature(Function Types/ User-defined Types) |
| [Zaremski1997] | Specification-based | Component Behavior | Specification Matching | |
| [Mili1997] | Specification-based | Functional Property / Component Feature | Functional Property / Component Feature | Specification |
| [Henninger1997] | Text-based / Keyword Indexing | Component Characteristics | Predefined Categories / User-defined Categories | Keyword |
| [Seacord1998] | Text-based | Component Interface | Component Interface | Content-based Keyword |
| [Damiani1999] | Context-based Description | Component Characteristics | Component Behavioral Characteristics | |
| [Inverardi2000] | Context-based | Component Behavior / Component Assumption | Component Behavior | |
| [Plasil2002] | Text-based | Component Interface | Component Behavior/ Component Interface | |
| [Chatzigeorgiou2003] | Graph-based | Message Exchange ( Component Function Call ) | Component Quality | |
| [Vitharana2003] | Knowledge-based | Component Structural Characteristics / Component Functionality / Business Rules/ Component Role in Use | Abstraction Hierarchy / Component Role in Use | The Use of Component |
| [Inoue2005] | Graph-based / Text-based | Component Association | Component Ranking / Content-based Similarity | Content-based Keyword |
It amazed me that the graph-based approach had been applied since 1992, although the approach in [Ostertag1992] was far from similar with recent researches.Most of them are functionality-based classification, only very few associated with non-funcitionall qualities. The knowledge-based approach seems to be an interesting direction based on these previous influential functionality-based approaches.
References
- [Seacord1998] R. C. Seacord, S. A. Hissam, and K. C. Wallnau, "Agora : A Search Engine for Software Components," IEEE Internet Computing, vol.6, no.2, pp.62-70, 1998
- [Chatzigeorgiou2003] A. Chatzigeorgiou, "Mathematical Assessment of Object-Oriented Design Quality," IEEE Transactions on Software Engineering, vol.29, No.11, pp.1050-1053, Nov. 2003
- [Plasil2002] F. Plasil and S. Visnovsky, "Behavior Protocols for Software Components," IEEE Transactions on Software Engineering, vol.28, No.11, pp.1056-1076, Nov. 2002
- [Inverardi2000] P. Inverardi, A. L. Wolf, and D. Yankelevich, "Static Checking of System Behaviors Using Derived Component Assumptions," ACM Transactions on Software Engineering and Methodology, vol.9, no.3, pp.239-272, July 2000
- [Zaremski1997] A. M. Zaremski and J. M. Wing, "Specification Matching of Software Components," ACM Transactions on Software Engineering and Methodology, vol.6, no.4, pp.333-369, Oct. 1997
- [Zaremski1995] A. M. Zaremski and J. M. Wing, "Signature Matching: A Tool for Using Software Libraries," ACM Transactions on Software Engineering and Methodology, vol.4, no.2, pp.146-170, April 1995
- [Mili1997] R. Mili, A. Mili, and R. T. Mittermeir, "Storing and Retrieving Software Components: A Refinement Based System," IEEE Transactions on Software Engineering, vol.23, No.7, pp.445-460, July 1997
- [Inoue2005] K. Inoue, R. Yokomori, T. Yamamoto, M. Matsushita, and S. Kusumoto, "Ranking Significance of Software Components based on Use Relations," IEEE Transactions on Software Engineering, vol.31, No.3, pp.213-225, March 2005
- [Damiani1999] E. Damiani, M. G. Fugini, and C. Bellettini, "Corrigenda: A Hierarchy-Aware Approach to Faceted Classification of Object-Oriented Components," ACM Transactions on Software Engineering and Methodology, vol.8, no.4, pp.425-472, Oct. 1999
- [Ostertag1992] E. Ostertag, J. Hendler, R. P. Díaz, and C. Braun, "Computing Similarity in a Reuse Library System: An AI-Based Approach," ACM Transactions on Software Engineering and Methodology, vol.1, no.3, pp.205-228, July 1992
- [Henninger1997] S. Henninger, "An Evolutionary Approach to Constructing Effective Software Reuse Repository," ACM Transactions on Software Engineering and Methodology, vol.6, no.2, pp.111-140, April 1997
- [Frakes1994] W. B. Frakes and T. P. Pole, "An Empirical Study of Representation Models for Reusable Software Components," IEEE Transactions on Software Engineering, vol.20, no.8, pp.617-630, August 1994
- [Vitharana2003] P. Vitharana, F. M. Zahedi, and H. Jain, "Knowledge-Based Repository Scheme for Storing and Retrieving Business Components: A Theoretical Design and Empirical Analysis," IEEE Transactions on Software Engineering, vol.29, no.7, pp.649-664, July 2003
下午4:09 | 標籤: CBSD, software reuse | 0 Comments