Background
Even as a developer you sometimes need to keep an eye on some things running on computer systems. If you are in a support or administration role it is even more important. Over the years I have developed or was involved in the creation of multiple systems in this ‘space’.
History
Years ago (14+ wow, feels like another life-time) I single-handily created and worked on a help desk tracking system called VHelpdesk first and then later renamed to Logidesk. That was back in my Vircom/Logical/3Fifteen days. At some point a version of this tool/system was even used by the Botswana government (via a service provider company). Those were interesting times… This system exposed me to the needs of service/operational oriented people that is/was helping other people in the IT world. You only really appreciate the needs of help/service desk people once you have to think like them.
In more recent years I’ve been more involved around the ‘BizTalk’ world – first for development but the last couple of years more on management and administration. In the process I had to create a lot of tools to help myself and others with these tasks. EventScavenger is one of those tools – but it is not the only one. Some of my tools are even used by others – even those in a full time production 24×7 environment. Yes, this is a little bragging – but if I don’t do it no one will 😉
There’s an art to it
Creating monitoring systems properly is sort of an art by itself – then again any decent computer system requires that. There is a fine balance between functionality, complexity and user experience. In short, your users define what the system must do. Never forget that! Sometimes you are lucky – if you are your own user you first hand know what you want. Similarly, if the users are people that does the same as you it is fairly easy to create that system.
Monitoring verses Alerting systems
Before I go on about monitoring systems lets first define the difference between monitoring and alerting – as I use them. Monitoring is purely the action of gathering information about the state of something and then record it so it can be viewed afterwards. Alerting is something that builds on-top of or extend monitoring to automatically alert users when some desired (or usually undesired) state is reached. As an example, EventScavenger is just a monitoring tool. It does not notify anyone or anything if it encounters some error logged by some machine’s event log. Microsoft SCOM is an alerting system (and more – it not only alerts but can actually automatically run some corrective script to attempt to fix something that has gone wrong).
Quickmon
A year or three ago I created something simple to monitor multiple types of computer system related entities with a plug-in architecture. I never really finished it properly but surprisingly it actually works to this day. Al it does is to gather some state data and present it as a simple icon on the taskbar (or originally in the XP day the notification area). You can open a detail window that display the state of each ‘agent instance’ to see what its individual state is. Further, you can also open an even more detail presentation of the particular agent’s details.
There are multiple agents – each with a different function but they all implement the same (.Net) ‘Interface’. Adding a new agent type is as easy as implementing the required interface and ‘registering’ the assemble/dll with the main app. The idea is/was that it should be extendable without having to change the main app at all. The agents I created so far (but never went on further) are: Ping, Service state, File count and one special one to monitor the number of suspended BizTalk instances (this was the one that gave me the idea). In the main app you can create a list of agents, mixing any number of them, duplicates of the same type of agent with different configs etc. This list of configs can be saved to a file (plain xml) so you can have multiple sets of things you want to monitor.
Agents
As an example: the Ping agent. This agent is configured to ping a list of machines – each entry has a maximum allowed ping time and time-out time. When the main app calls this agent instance it only returns one value – the ‘worst’ state of all the states of all the ‘pings’. If (any) one ‘ping’ times out the total state is ‘error’. If any one ping exceed the maximum allowed value the total state is ‘Warning’. Otherwise the total state is ‘Good’ or OK. You get the idea. How/what an agent define as error, warning or good is up to the implementation of that particular agent.
Other than the main GetState() method the agent must also support (Windows Form) interfaces for (1) ‘editing’ its config and (2) displaying details about the states of things it monitor.
What has become of Quickmon
Quickmon sounds all nice but at some point the complexity and time (and willpower) to go on stopped me to develop it further. As it is it works like explained above but I have some new and additional ideas I would like to add. The ‘current’ version is purely visual monitoring tool. It does not record history. I never completed an easy way to initially set up the app with its registered agents (have to be done manually so far). It was created in .Net 2.0.
Funny thing is, I started using it again ‘as is’ but I decided to start a new project based on the same idea and extending it. The only issue (again) is time.
New project
The new project will take the old Quickmon concepts like a plug-in architecture, grouping or rolling up resulting states and having each agent define its own functionality. Some new ideas I have are:
- Adding alerts/notifications – also using a plug-in architecture. e.g. Log file, RSS feed, SMTP, database table.
- Make it into a Windows service (but also still have a UI part)
- Having dependencies between agents – e.g. If a machine is not pingable then don’t bother checking any service states…
- Adding some more agents – e.g. Disk space, Database/table size, Event logs (touching on EventSavenger grounds here) etc.
One problem with all these ideas is that the more functionality I add the more complex it becomes to manage – this is true for any project. Also, it might start competing on territory where commercial products are – like SCOM… and I don’t want that…
I have started with this project and have some working ‘bits’ already. Perhaps I must see if there are others that might be interested in collaborating on this project – just an idea. Any takers?
0 Comments.