Saturday, December 15, 2007

Spock.com: Identity Manager

You ever wonder if there are sites which manage people information which is found on the web. It is an interesting concept. If you type my name, Enoch Moses, in Google.com, you get Christian theological sites which discuss about Enoch and Elijah. You also get my entry in Linkedin.com. However Google displays a number of my blog entries, my linkedin profile, myspace page, etc., and etc. Sites like Spock.com let you manage your information on the web and this way you can define and identify which is your content, content about you or content involving you. It's a neat site. Here is my Spock page.

Enjoy!

Wednesday, December 12, 2007

Would I invest in a dot com startup?

Every Joe on the street thinks he can make a quick million or two by starting his own dot com company. Hey Chad Hurley and Steve Chen, founders of YouTube.com did it. Unfortunately there is a saturation of the web sites which promise functionality however they are dependent on your, the user's, data. The first question to ask is why would I want to put my information on a third party web site so that web site founder can make some money. It is the sad truth but most of these companies will not last long. For example let's look at various Social Networks:
  • MySpace.com - one of the first site which actually picked up in popularity. Now it is mired in mediocrity and I don't see anything new and exciting happening on the site. Did I mention that they have added Facebook like functionality.
  • Facebook.com - I have to say that I was a skeptic when I joined this site however this site offers neat functionality where the user can actually spend time on the site. I am a big fan of Facebook's scrabble application
  • High 5 - Started by an east indian and it's being marketed heavily in the east indian community.
  • Orkut - I have to say this is probably one of the worst social networks I have come across. It is probably Google's worst purchase.
  • Kadoo - A new social network site whose UI looks promising however I still don't have an incentive to join this group. I am not into propogating my identity across the internet
  • Linkedin.com - I like this website. It's a social network for your professional contacts
  • Xing.com - It is a similiar network like Linkedin however this one is popular in Europe
  • PageFlakes.com - Someone emailed me asking me to experience this site. Once again I ask the question. Why should I sign up on PageFlakes.com?
  • NING.com - This is the UBER social network where anyone can create their own social network. UI is not that great and it is meant as a research application.
After I mentioned all of these social networks, how would a investor invest in these kind of businesses? Frankly everyone of these social networks offer similiar if not identical functionaliy. I would probably want to invest in networks which have alot of users and the network has a niche like Linkedin and Xing which only work with professional social networks.

Friday, December 7, 2007

Items of poor design...

For the last six months or so, I have been in involved in a project where is the system has been growing organically. It has been growing like fungus. Fungus is a unique living organism. It does not have a head, legs, hands or a body. It just exists and it keeps evolving as long as there are enough nutrients and dampness. The system I have been working on is an IT fungus. The system has evolved over the last four years or so. Developers come and go but this system still exists. This system has no requirements documents, design documents, and no test plans. The software is poorly documented and I wonder how this system still exists. Well it does exist and it seems to be growing larger and larger. The system is composed of subsystems which have evidence that developers have tried to improve the system but they have miserably failed. The failed implementations were not cleaned and they tend to future the process this system's evolution. This system is a J2EE system and you will notice the following things in this system:
  • 1500 to 7000 line JSP page which have scriplets embedded in javascript which inturn invokes JDBC calls
  • Partially implemented hibernate framework. This is evident with *.hbm.xml files in the source code
  • Numerous properties files which are now neatly packaged in a oracle database
  • Spring Web Flow - The developer who implemented this portion did a great job.
  • Prototype Ajax
  • JSP pages with scriptlets and jstls
  • Same piece of code in numerous web apps which have been customized for each web app
  • Field level filtering in the database for each user (not role but user)
  • User is authenticated between each web-app even though each web-app is part of the bigger system.
  • Custom API integration for each COTS product.
  • Outdated Stored Procedures in the database which were not used anymore.
A few weeks ago, I wanted to write junit unit-tests for some of the "cleaned up" classes and I failed miserably since all of the code is tightly coupled and it follows the onion architecture. In an onion architecture, you never know what you are going to get under each peel. The point of this blog entry is to remind every IT personnel that systems are not organic beings but they are rather simple business logic processors. In the SDLC, a good design, which has been reviewed and analyzed, clean implementation and robust test cases are the basic ingredients for a good system.
(PIC of a fungus called WitchButter)

Monday, December 3, 2007

Book Review for Mike Daconta's new book

Couple months ago, I agreed to review my esteemed colleague Michael (Mike) Daconta's new book Information As Product. I liked Mike Daconta's previous book called The Semantic Web. The Semantic Web was a great book and it basically laid out the history of the semantic web, described the benefits of semantic Web and described the vision of the semantic Web. This book was my passport into the world of ontologies, data modeling and understanding the importance of data. I have also had the privilege of working with Mr. Daconta on couple of DHS projects and I believe he is truly a visionary in data management.

The biggest knocks against Mike Daconta in the industry is that he is a "dreamer" and he has not been able to deliver his dreams into substance. I believe Mike is a visionary and people like him are essential in IT innovation. He offers ideas which address actual business problems and it is up to engineers to formulate the ideas into reality.

His new book Information As Product follows this pattern. In the book, Mike offers a solution of producing Information as a product which comes out of a "Information" factory line and it is appropriate information for the appropriate person and it is delivered in the appropriate time. The book presents a general solution to a problem plaguing various enterprises. The problem is that there is a temporal and semantic gap between information consumers and producers. The book does a great job of describing the concepts involved with this idea however no system analyst and architect can decompose this book into functional and non-functional requirements to build a system which will make concepts in this book a reality. Mike states in his book that the book Information As Product is the first book in a series of books which will engage its readers in a dialogue on how information systems can be improved.

The things I liked in the book are:
  • easy to read
  • presenting the ideal information management system as a factory line where information can be packaged in a package
  • I loved the way he describes packaging up the information. I believe this idea can implemented
I wish he expanded these concepts better in this book:
  • the importance of metadata, selecting the right metadata and consequences of poor metadata management
  • The DIKW (Data-Information-Knowledge-Wisdom) pyramid - It is only a conceptual model. I would love to see how metadata fits in this pyramid.
  • He lost me in couple of parts otherwise it is not a bad book.
In summary, if you want to build a system from this book then I recommend that you don't buy this book. If, on the other hand, you are looking at data management solutions for your enterprise then this is a great book since enterprise level functional requirements can be derived from this book. I, personally, enjoyed reading the book however I was left with more questions than answers.
Links to buy the Information As Product book.

Thursday, November 29, 2007

Day 2 at the Ontology Conference

My second day at the ontology conference was quite good. I impressed by the various teams which presented their papers on they were designing and implementing ontology based systems. The biggest theme from the conference was that the current technology and lack of defined ontology methodologies was the biggest drawback in this field. Most of the applications are prototypes and they are extremely slow in processing decent sized ontology. I heard talks about reasoners, owl, geo-spatial ontologies, multi-order logic processors, ontologies in graph databases, etc., etc. However the applications which were using these technologies were prototypes. The other problem was that if ontologies were not built correctly then the results were hideously wrong. Everyone in the conference agreed that ontologies and their applications are still new in the field of IT however the promise of ontologies and their applications is so great that large organizations keep funding R&D in the field. I personally enjoyed my time at the conference since it was good to see other data lovers and people who understood the value of data in any IT enterprise. I will probably go again next year. Here is a list of products which were mentioned in the seminar.
  • Knoodl.com - A semantic wiki. It creates ontologies from the wiki entries or uses uploaded ontologies in categorizing wiki entries.
  • VideoQuest - This product searches entities in a video. For example, if the user typed in the query "white car in saint louis" then the result set would include videos which have a white car in Saint Louis. The backend of this product is based off ontologies.
  • Poised For Learning - Rensselaer Polytechnic University's Rensselaer Artificial Intelligence and Reasoning (RAIR) Laboratory's Ontology product which is based reasoners. I was very impressed with this research.
As you can see there weren't many products since this area is still new. That's all for now.

Wednesday, November 28, 2007

Day 1 at the Ontology Conference

Today I went to an Ontology conference which was sponsored by the National Center for Ontological Research (NCOR). I heard talks from various vendors, implementors and subject matter experts in the field of Ontology. Before I get into what they talked about, let me first state the definition of what is an ontology. Wikipedia defines an ontology as, "... a study of conceptions of reality and the nature of being." What does that mean??!! Well it is basically an exercise where ontologists (people who create and work with ontologies) model existing systems into categories which then could be used by information systems to make inference relationship (relationships which are not obivious to the human). As you see it is a qualitative field, people will argue how data entities should be modeled so that they can accurately and precisely describe an existing entity. For example, Ontologist A says that every plane needs to have pilot while Ontologist B would say this is not true. He would say that there are drones which are small planes, that are controlled by a computer. One of them may be right or both of them can be right. The key for any ontology is that it needs to be in a domain to avoid confusion. So in theory both ontologist A and B are correct if it is assumed that both of them work in different domains. Ontologist A is creating an ontology for the airline industry while Ontologist B is creating an ontology for some country's military. Issues like this cause massive headaches to any organization's Chief Information Officer (CIO) since the data cannot easily be exchanged across organizations and across domains. If Ontologies are created well then they are quite powerful.
Currently the ontology is wonderful in the realm of philosophy and theoretical computer sciences. however there isn't much technology in the field. I have been off and on working with ontologies and they are not the easiest thing to implement. I, however, am a strong believer that a good ontology can be used to validate data architectures which pertain to the ontology's domain.

At the conference, I heard two great presentations on ontology. They are:
  1. Werner Ceusters - He gave a fascinating talk on what is ontology which incorporated Basic Formal Ontology (BFO).
  2. Steven Robertshaw - He gave a talk on the lessons learned in implementing complex ontologies.
These people understood how to use ontologies and how not to use ontologies. Both stated that the technology is lacking in this field. Werner and his team built an open source project which lets you play with ontologies. The project, which is a Java based application, can be found at http://sourceforge.net/projects/rtsystem. To get a better understanding on how ontologies work, I recommend downloading the Protege product. It, too, is an open source product maintained by Stanford University. Protege lets you build data models but at the same time it lets you validate your model with "instance" data. I got fascinated with field couple years ago when I was asked to validate a data model for one of the US government agencies. If you are totally lost about what I am talking about then please check out the following links:
The day one of the conference also consisted of vendors displaying their products which are based on the principles of ontologies. In my next blog entry, I will give a more detail summary on what happened at this conference. Stay tuned! (The image is a visual representation on an ontology)

Are RESTful web services ideal in a SOA?

Representational State Transfer (RESTful) web services are the latest buzz in the web services domain. What are RESTful web services? RESTful web services consists of HTTP clients which send requests as request parameters in the URL and the response is a XML document which can be viewed in a browser. This is ideal service for technologies like Asychronous Javascript and XML (AJAX) which take XML responses and apply XML transforms (XSLT) to generate a visually pleasing display of content. AJAX's primary mode of transport is through the HTTP browser and it's mode of transport is HTTP or HTTPS. Currently other developers and vendors like Yahoo!, Google and Microsoft have harnessed the AJAX technologies, RESTful services and RSS feeds to generate data aggregators which allow users to find, aggregator, and filter from various data sources how ever they are merely point-to-point services. AJAX working with RESTful web services is the ideal model for the federated query where one request is federated across multiple data sources and then the responses from each data source is aggregated to show the "full" picture.

Yes RESTful web services are not that bulky since they have don't have the SOAP wrapper however let us not get caught up in RESTful services are ideal in a SOA enterprise. The problem with REST is that it is designed for point-to-point services with minimal reuse. This is the same type of problem where SOAP based services exposed method calls which required specific datatypes in their Web Services Description Language (WSDL). When it is a point-to-point services, alot of work is spent in formatting, transformation and processing information to fit the specific datatypes in the non-reusable methods. To make a SOAP based web service more SOA friendly, it is recommended that the services take XML documents as the parameters in the web services method calls. This way one XML document can be passed around where the contents of the XML document might be altered by various method calls. It is not practical for RESTful web services to pass a reference to a XML document as a request parameter.

RESTful web services are great with AJAX however I wouldn't recommend using RESTful web services for services which are highly popular, reliable and have a potential to be reused alot.