Wednesday, May 18, 2011

Ruby on Rails vs Java ORM

Wow, I am having fun with Ruby on Rails. The migrations and scaffolding alone make me less of a coder and more of an Engineer. Lol, actually, it is kind of like that movie Wall-E. You know... where all the humans relied on the robots to do everything, evolving the humans into weak boned blobs. Rails does a lot for you, and it is hard for a Java programmer, like myself, to give up the deep control I had over the software. Or is it???

One of the concepts I learned today was the use of named scopes in Rails. The concept is that you can predefine a db query, which can be dynamically added to any normal method call on the Model. For example, lets say you have 2 tables, Customer, and Address. A Customer can have one or many Address records. So, I want to capture that in my Customer model. The default functionality for the Customer model would be to return one, or many, Customer records, using the find() method. Which is good, but I also want to return the Address information as well. Consider the following Customer model:
class Customer < ActiveRecord::Base
has_many :addresses, :dependent > :delete_all, :foreign_key => 'person_id'
named_scope :with_address_info, {
:select => "customers.*, addresses.*",
:joins => "JOIN addresses on addresses.person_id = customers.id",
:order => "customers.last_name asc" }
end
The above code is all I need for a simple Customer model, customer.rb. The has_many identifier tells Ruby that the Customer can have many Address records. The "with_address_info" named scope allows me to dynamically include the Address records with any query. In my Controller, I can have the following code:
customers = Customer.find(:all)
customerFull = Customer.with_address_info.find(:all)
That's all. Pretty powerful stuff. That small model file has given me the ability to Create, Retrieve, Update, and Delete (CRUD), all in 7 lines of code. Technically, if I don't care about the named scope, the size of the file would be 3 lines of code. But that seems pretty scary. It is like magic. Obviously, it isn't magic. The Rails framework writes all of those extra lines of code for you behind the scenes.

So, does it really take away control from a programmer, more than Java? Well, if you are use to manually creating JDBC calls to a database, then YES. You have a lot more control with the DB access in the code itself when you work with the straight JDBC drivers.
But, the need for that level of control is rare. Many of us use Object Relational Mapping (ORM), like Hibernate or Java Persistence API (JPA). In that case, Java takes the JDBC overhead away as well as the Rails framework. Of course, the syntax between Ruby and Java is such that a Ruby Model is much smaller in lines of code then a JPA entity (all those getter and setters add up). There is also the dynamic typing that I won't get into. But if you consider the "magic" with mapping a software object to a database object, then they are performing the same actions. Consider the following JPA Entity:

import java.io.Serializable;
import javax.persistence.*;
import java.util.Set;
@Entity
@Table(name="customers")
public class Customer implements Serializable {
@Id
private int id;
@Column(name="customer_type")
private String customerType;
@Column(name="email_address")
private String emailAddress;
@Column(name="first_name")
private String firstName;
@Column(name="last_name")
private String lastName;
@OneToMany(mappedBy="customer")
private Set<Address> addresses;
public Customer() {
}
public int getId() {
return this.id;
}
public void setId(int id) {
this.id = id;
}
public String getCustomerType() {
return this.customerType;
}
public void setCustomerType(String customerType) {
this.customerType = customerType;
}
public String getEmailAddress() {
return this.emailAddress;
}
public void setEmailAddress(String emailAddress) {
this.emailAddress = emailAddress;
}
public String getFirstName() {
return this.firstName;
}
public void setFirstName(String firstName) {
this.firstName = firstName;
}
public String getLastName() {
return this.lastName;
}
public void setLastName(String lastName) {
this.lastName = lastName;
}
public Set<Address> getAddresses() {
return this.addresses;
}
public void setAddresses(Set<Address> addresses) {
this.addresses = addresses;
}
}
The Customer entity is a lot more wordy then the Customer Rails Model, but it performs the same kind of magic, just a little differently. If you were to use this Entity in a session EJB, you could have the following code:
List customers = entityManager.createQuery(" from Customers c ");
This code already has a reference to the address records. You can just call the getAddresses() method:
for (Customer customer : customers) {
List<Address> addresses = customer.getAddresses();
....
}
So, overall, Ruby on Rails provides a quick, and less verbose method of getting data across multiple tables. But if you use some popular ORM packages, you really arn't losing the control you thought you had.

Saturday, January 01, 2011

Happy New Year

Well, 2010 is gone. I liked 2010. I've learned a lot and expanded into a new career. My new years resolution is to keep on learning, and keep up on the blog. Blogging is fun, but it is also a great way for me to keep my skills up. So, happy new year to everyone.  Keep coding.

Thursday, December 30, 2010

Java EE Packaging

During my off hours, I do a little teaching. The course I teach relates to Java EE and EJB development. Over the past 3 years that I've been teaching this course, I've realized there are quite a few common issues that students get hung up on. Initially, the issues with classpaths, packaging, and deployment are the biggest hurdles. So, here is an overview about the packaging of Java components. I just JBoss as an example, but most Java EE servers are handled in a similar way.

First, we have the Java ARchive, or jar, file. A jar file is one of the earliest package formats in Java, and can be used to compress, and package classes together in a neat library. In the EJB world, the jar is what you use to package your EJBs. You can package EJBs by themselves, and deploy them separately, or they can be included as an Enterprise application (more on that later). The jar file contains your classes, and a META-INF directory. The META-INF directory contains config and deployment descriptor files. To create a JAR file, just go to the top level of your class hierarchy and type "jar cvf jarFile.jar .". This will package your classes, properties, config, and deployment descriptor files into a single jar file. From here, you can deploy it straight to the Java EE server. It should be noted that some Java EE servers require some extra deployment descriptor files, refer to the vendor documentation for more information. However, in the case of JBoss, you can just copy the file into the "deploy" directory, and JBoss will run with it.

Upon deployment, JBoss will unpack the code, and set up the EJBs to be used. During the deployment, errors can happen. This is a good time to pay attention to the logs. Many problems happen with the deployment descriptors, or when you are using injection, with annotations. So, looking at the logs will help you to determine if the problems are runtime, or deployment time. A successful deployment would contain something similar to the following in your server.log, or standard output.
2010-12-29 10:01:08,666 INFO [org.jboss.ejb3.session.SessionSpecContainer] (HDScanner) Starting jar=MyTestEjbs.jar,name=HelloWorld,service=EJB3
2010-12-29 10:01:08,677 INFO [org.jboss.ejb3.EJBContainer] (HDScanner) STARTED EJB: awg.ejb.session.HelloWorldImpl ejbName: HelloWorld
2010-12-29 10:01:08,752 INFO [org.jboss.ejb3.proxy.impl.jndiregistrar.JndiSessionRegistrarBase] (HDScanner) Binding the following Entries in Global JNDI:
HelloWorld/remote - EJB3.x Default Remote Business Interface
HelloWorld/remote-awg.ejb.session.HelloWorldRemote - EJB3.x Remote Business Interface
The above log entry not only tells me that the ejb was deployed without problems, but it also give me some information about the ejb. For example, the JNDI name. In this case, the JNDI name is "HelloWorld/remote". This is very useful for when you are writing your client code.

Once the EJB is deployed successfully, you can use it in some client code. You can write a local client, or you can create a web based client. The important thing to remember when writing a client to access an ejb is that ejbs are considered remote objects. So, you need to supply some information in your code on where to find the references to the ejbs. This is commonly performed using the JNDI. If you are accessing the ejbs from a jsp, and it is deployed on the same server instance, then the access to the ejb is pretty simple. Just get a reference to the Initial Context of the JNDI, and pass in a String, for the name of the resource you are trying to get. For example, if we want a reference to the HelloWorld ejb I deployed, I would use the following code:
InitialContext ctx = new InitialContext();
awg.ejb.session.HelloWorldRemote hwr = (awg.ejb.session.HelloWorldRemote) ctx.lookup("HelloWorld/remote");
However, if we are using a client remotely, we need to write code that is more involved. For example, consider the following command line client:
package awg.client;

import java.util.Properties;
import javax.naming.Context;
import javax.naming.InitialContext;
import awg.ejb.session.HelloWorldRemote;

public class CommandLineHelloEJBClient {

 public static void main(String[] args) {
   try {
     // create an InitialContext, with JNDI properties
     Properties p = new Properties();
     p.put(Context.INITIAL_CONTEXT_FACTORY, "org.jnp.interfaces.NamingContextFactory");
     p.put(Context.PROVIDER_URL, "jnp://localhost:1099");
     Context ic = new InitialContext(p);

     // Get the a reference to the Remote Object from JNDI
     HelloWorldRemote client = (HelloWorldRemote)ic.lookup("HelloWorld/remote");

     // run methods on the ejb
     client.....();

   } catch (Exception e) {
     e.printStackTrace();
   }
  }
}
In the above example, I am giving the URL, jnp://localhost:1099, and the Context type.

With local clients, we can wrap them in jar files as well and run them from the local command line. But what about web resources, like JSPs? In the Java world, web resources can be packaged up in their own package as well. The Web ARchive, or war file. The WAR follows a specific directory structure to allow for easy deployment of web resources.

The top level directory is considered the webroot, and it contains JSPs, HTML, CSS, JavaScript files, images, and any other file you want to be exposed to the web. Within the webroot, there is also a WEB-INF directory. Inside the WEB-INF directory, config files, and deployment descriptors, are kept. Also in the WEB-INF directory, there is a directory for your classes, and a lib directory, for any libraries your'd like to package with your web application. To create a war file, just go to the webroot directory and type "jar cvf warFile.war .". This will package your web files, and additional classes and jars into a war file. From here, you can deploy it your Java EE server, just like you did with the ejb jar file.

The jar and war files are technically all you need to run your web files, and ejbs. But to make the deploying and submitting even easier, and cleaner, you can use an Enterprise ARchive, or ear, file.
The ear file is nothing more than a single package containing your war file(s), jar file(s), and deployment descriptors. To build an ear file, follow these steps:
  1. Package up your ejbs into a jar file
  2. Package up your web resources into a war file
  3. Throw your war and jar files into a directory where you can package them together
  4. Create a new directory called META-INF (under the directory containing your packages)
  5. In the META-INF directory, create a deployment descriptor called application.xml.
    * The application.xml file defines the modules in your application.
  6. From the directory containing your packages, package everything up as an ear file: jar cvf MyApplication.ear .
  7. Deploy the ear to the JBoss deploy directory.
Seems like a lot of work, however, you can build Ant script, or use an IDE like Eclipse or Net Beans to do all of the heavy lifting. I'll talk more about that in another post.

The ear file makes deployment much more organized, and testing of the module is simpler. One thing to remember is that the Context lookup will be different with an ear, compared to using a war and jar separately. The ear context is basically the name of your ear file. For your lookup, you need to reference the ear context, when using JBoss. For example:
  • I created an ear called MyApplication.ear. The file contains all of the examples from the class. changed all of the lookups to include "MyApplication/" at the beginning of the look up. If I deployed the war and jar files separate, the code in the HelloWorld client would look like this:
    HelloWorldRemote hwr = (HelloWorldRemote) ctx.lookup("HelloWorld/remote");
  • However, since I am deploying it as an ear file, the code now references the ear context:
    HelloWorldRemote hwr = (HelloWorldRemote) ctx.lookup("MyApplication/HelloWorld/remote");
It should be noted that the JNDI names are not standard. Each Java EE server has it's own way for referencing ejbs on the JNDI. You may need to refer to documentation to get the format of the name. You can also use the ejb-jar.xml file in the META-INF directory of you ejb package, to customize the name, but there are still vendor specific formats tied in. That is why they are called by string names. If you write your code correctly, you can hold the names in a properties file, XML, or even a database. This will keep your code more portable. However, in EJB 3.1, JNDI names will be global and will follow a standard pattern that will make your code more portable across different application servers. But that is another story.

For more information, you can refer to the JBoss "Getting Started" documentation: https://www.jboss.org/file-access/default/members/jbossas/freezone/docs/Getting_Started_Guide/beta500/html-single/index.html#Configuration_Files

Monday, December 27, 2010

Installing & Running JBoss in Ubuntu

Being mostly a Java guy, a couple of my favorite servers to work with are Tomcat and JBoss. While at work, it has been exclusively Tomcat, lately, I've been mostly interested in working with JBoss. For a couple years, I've had JBoss installed on my Windows machine. It is pretty easy,just unzip on your hard drive and execute the run script in the bin directory.

Since I installed Ubuntu, I've wanted to switch to using Linux as my JBoss OS. With Ubuntu, I could easily use the update manager, or apt-get jboss, to get it installed. However, it wouldn't give me the version I wanted. So, I figured I'd work this manually.

The first thing I do, is download a zip. A common place to put the software is /usr/local. So, I will download, and unzip the package there.
cd /usr/local
sudo wget http://sourceforge.net/projects/jboss/files/JBoss/JBoss-6.0.0.Final/jboss-as-distribution-6.0.0.Final.zip
sudo unzip jboss-as-distribution-6.0.0.Final.zip
The funny thing is that I can't seem to find a tarball for JBoss, so I just get the zip. Based on this, and the directory structure of JBoss, it seems that the development of JBoss was mostly geared towards Windows based servers. Regardless, it works great on Linux as well.
Now, at this point, you can just run JBoss out of the box, using the run.sh script in th ebin directory.
sudo /usr/local/jboss-6.0.0.Final/bin/run.sh
This will start up the server, and the terminal window would be the standard output. It that is all you need, then you can stop right here. But in general, there are a couple things to simplify the start up and shut down process. I like to control the server as a service, from the /etc/init.d directory. Within the bin directory, there is a startup script which will call the run.sh script as the jboss user. The script is written with redhat in mind, but is easily modified to run with Ubuntu. But first, lets create a jboss user and grant ownership to the jboss directory.
sudo useradd -s /bin/bash -d /home/jboss -p jboss
sudo chown -R jboss:jboss /usr/local/jboss-6.0.0.Final/
The useradd command will create the jboss user, with the /home/jboss directory as it's home, and the default shell is bash. This command will create the jboss user without a password. If you are like me, and like to make things as secure as possible, use sudo passwd jboss to create a password for the jboss user.

sudo ln -s jboss-6.0.0.Final/ jboss. The chown command grants ownership to the jboss user. The -R, after the chown command, does a recursive call to the sub directories.
Now that the jboss user is set up, copy the jboss_init_redhat.sh file into the /ect/init.d directory, and name it jboss.
sudo cp jboss-6.0.0.Final/bin/jboss_init_redhat.sh /etc/init.d/jboss
There are a couple changes we need to make with the file. first, it assumes that the JBoss home directory is /usr/local/jboss. By default, our command to unzip the package created the home as jboss-6.0.0.Final/, so we can either rename the directory, with the mv command, change the init file to reflect the real path, or create a symbolic link. I prefer the symbolic link, because we can keep the old packages as we download more, and if we need to revert, it is just a matter of modifying the symlink.
sudo ln -s jboss-6.0.0.Final ./jboss
The next change is in the file itself. Use VI and modify the JAVAPTH variable to the correct path of your java command. In my case, and most likely yours as well, the path is /usr/local/bin. If you don't know the path, just use the command "which java". Verify some of your other variables, just in case you need to make changes for your system.

At this point, you are all set to run it as a service. Use the following commands to start, stop, and restart the server.
/etc/init.d/jboss start
/etc/init.d/jboss stop
/etc/init.d/jboss restart
If you created your jboss user with a password, you will need to supply the password. The server will start up and the logs will be written to /usr/local/jboss-6.0.0.Final/server/default/log/server.log. If you'd like to have the server start at system start, then you can update the rc.d to include it as an auto start up.
sudo update-rc.d jboss defaults
I personally prefer to start it up as needed. Once started, you can deploy your jars, wars, and ears as you normally would, in the deploy directory of the server you are using. In this case, we are using the default server.

Update 12/29: I updated this post to reflect the GA release of JBoss 6.0.

Wednesday, December 22, 2010

Ubuntu Wubi vs Grub Upgrade

Back in my younger days, I really loved tinkering around with Linux. I wouldn't consider myself a Linux expert, but I was learning, and gained a lot of experience. I stated out using Red Hat, before Fedora, and played around with Gentoo, PHLAK, and even built a firewall using Coyote Linux installed on a Floppy disk. It was a great learning experience, and I made sure I used a computer which I didn't need anymore.

Then, I had kids. All of my free time was taken up by the family. I needed to prioritize my life a little stricter, and family, personal health, job, and relationships took priority over Linux. So, for the following 5 years, I was a Windows user. My job didn't require any Unix or Linux work, and being a Java programmer (for the most part), it didn't matter what OS I used.

Last summer, I decided to get back into it. My kids are grown to the age where they can entertain themselves, and don't need constant supervision. So, I decided to check out Ubuntu. I've heard a lot of great things about Ubuntu. I went to their website, and was about to create a LiveCD when I noticed they have a Windows Installer.... huh? Whats this? An executable which installs Linux as a Dual Boot? No hard drive partitions? No figuring out swap size? The culprit is known as Wubi. after a download and a double click, I restarted. My computer started up with a Bootloader menu, displaying a selection for Windows 7 and Ubuntu. I selected Ubuntu and set it up. It was so easy. I almost felt ripped off. I mean, come on, my elitist Linux user attitude could no longer be expressed, since Ubuntu is as simple as Mac OS.

All was going well, for a couple weeks, I was in dual boot bliss. Until, one day, the Update Manager notified me that there was a new version of Grub. I refreshed my memory about what Grub was. Grub is a bootloader. Ok, I will update it.

Then it happened.... a problem with the typical Linux complexity. I rebooted, and in place of the nice bootloader menu, I got the following:
error: no such device: c73d2390-fb23-82c2-63823ae2eac2c
grub rescue>
What is this???? I quickly grab a second laptop and start up google. I search for "grub" and "ubuntu grub rescue". A lot of results are returned. After a while of sifting through the links, I find out that Grub and Wubi don't really play nice. In theory, Wubi has a bootloader which should send control to Grub. That is what was happening before the upgrade. The new version of Grub caused a conflict, such that Wubi didn't run, and Grub didn't know where it was. My initial google search was good for learning about the problem, but didn't yield much in terms of fixing the problem. So, I added "wubi" to the search terms, and I received a lot more results. Many of them told me to go here: "How to restore the Ubuntu/XP/Vista/7 bootloader (Updated for Ubuntu 10.10)". This seemed like the fix I was looking for. Only problem, it wasn't ment as a fix for Wubi installs. Wubi is a different beast all together.

More searching of the ubuntuforums site produced more insight about the problem, and how to ultimately fix it. The solution was to install lilo, another Windows bootloader, and don't ever update Grub. So, I made an Ubuntu LiveCD, and booted from it. That worked nicely and gave me access to a terminal. I then installed lilo, and added it to the mbr (the first sector of the primary hard drive):
sudo apt-get install lilo
sudo lilo -M /dev/sda mbr
After reloading.... It worked :-)
Other forum posts recommended to lock the Grub version in the updater. That is ok to do, but I decided I don't like Wubi anymore. But what can I do? I have a lot of customizations and software currently installed in my current Ubuntu install.

Well, it turns out that if I had only installed Ubuntu as its own partition from the beginning, this wouldn't have happened. But I did not give up hope, googleing "convert wubi install to partition". I found this post: "HOWTO: migrate wubi install to partition". Perfect, just what I needed. Basically, I have to free up some space on my hard drive, aka deallocate it, and create 2 partitions. One for the root /, and the other for the swap space. Windows 7 has an easy way to "shirnk" a partition, which will free up some space on a hard drive. I made the 2 partitions with 60G for root, and 6G for swap (I have 4G of RAM, so 6G sounded good).

Going through the manual steps outlined in the post worked like a charm. I rebooted and could use an updated Grub to load into Ubuntu or Windows. I booted into windows and used "Add/Remove Programs" to delete the Wubi install, since I no longer needed it. All is right in my Ubuntu world. The moral of the story is.... learn to use google.... and beware of Wubi. Wubi isn't bad, but it wasn't obvious that this would be a problem. After all of this.... my elitist Linux user attitude is restored (At least in front of my non-Linux friends).

Wednesday, December 08, 2010

You Know One Programming Language, You Know Them All…(sort of, don’t forget…. Time is Money)

I know that all two of my faithful readers out there (hi mom, hi dad), want to know how my job search has gone. Well… it has gone really well. I found myself a really good gig working for a product based company, that writes software for network tools. But make no mistake, the interviews I went on made me feel like I was back in college. All of a sudden, I was up all night studying the different algorithms to sort trees, explain TCP/IP, and making sure I can discuss the languages I know like the back of my hand, and at least show some understanding of the languages I don’t know.

The job who offered me the job had a pretty grueling, 5 hour interview. I guess I did pretty good considering that many of the questions revolved around Ruby, Perl, and the Rails framework. All of which are concepts I’ve only read about and briefly touched in my early days. During the last 2 weeks at my previous job, I got into some debate about the importance of Computer Science and Software Engineering fundamentals. My old job consisted of quite a few people who believed that if you are a pretty good programmer in one language, then you can easily move to another language. They believe this because, other then some syntax differences, all languages are the same… Right?  Well, sort of.

You see, some languages are compiled, some aren’t. Some languages are object oriented, some aren’t. Some languages are interpreted, and some aren’t. These are all important attributes of a language, and will affect how your code. Interpreted languages are more flexible, but can have some performance issues, so you need to keep that in mind if speed is a need. Object oriented languages are wonderful to design with, but tend to carry a lot of overhead. What if you are a Java programmer, taking a job writing C code? That’s when you realize that garbage collection has made you dependent. But that’s not all. Different software has a different purpose, which is important. Are you going to be writing client/server code? Client only code? Web based applications?  The concept of using HTTP is much different then writing a client side application. 

So, where am I going with this?  I managed to convince this company, which works in an area that I am unfamiliar with, with code that I rarely, if ever use, to hire me. It wasn’t only because I am a pretty good Java programmer, but I showed that I understand computers, networks, software, and the fundamentals of Computer Science and Software Engineering. I showed that I have the initiative to learn what I don’t know, and take the challenges head on.If I would have went into the interview with just my Java history, then I wouldn’t have gotten the job. What my old coworker didn’t understand is that time is money. If you cannot prove that you will be up to speed in a timely manner, then they have to make their decision based on that. Companies don’t want to waste time and money on someone who doesn’t understand fundamentals. Especially when there are others out there who want the job too. You only have a short time during the interview to prove why you are better then then next guy.

During my first couple week at the new job, I’ve been learning all I can and diving right into Ruby on Rails with ext JS. My oop and Java history has helped me pick it up pretty quickly, but I made sure going in that I already started learning. There was one mistake I made the other day, that I spent hours trying to figure out. It wasn’t until I asked the guy in the cube next to me what I was doing wrong, and he picked out a syntax mistake right away. I didn’t realize that back ticks ( ` ) actually meant something Smile But it did show me that syntax is not as trivial as many believe. So my advice to you is to learn about the company that you are interested in. And learn their technologies. At least be able to discuss the concepts, but don’t be afraid to tell them you don’t know something. Just as long as you let them know that you will know it.

Tuesday, November 09, 2010

My Common Interview Questions… Technical and Introspective

With all this focus I’ve had on finding another job, I started to look at the questions, and standards I have when interviewing a candidate in my company.  My upper management doesn’t like my questions, because I tend to lean more towards the computer science/software engineering side of the candidates skill set. Where my management would rather me evaluate their ability to show up to work on time (not that I’m bitter or anything). You see, I work for a service based company. One that get’s paid for “warm bodies”. They don’t have a large amount of pride and discipline for the engineering craft. Which is one of the main reasons I am looking for another job.

Anyway, I digress, my focus on what I bring to an interview, has been improved by my own interview questions. I like to use these to determine if someone is as average a web geek as I am.  I've collected a bunch of questions that are technical, but also inquire about the candidate’s professional and technical personality.

General Questions:

  • What sort of websites, blogs, and/or user forums do you follow? This is an open question, and there is no right answer. But anyone who is claiming to be technical should have a decent list of online resources for which they can learn, and stay on top of topics. For example, on a daily basis, I browse Slashdot, darknet, coding horror, DZone, Ben Nadel's ColdFusion blog, OWASP.org, and javablogs.com, among others.
  • If you are developing an application, and you come across an error that you've never seen before, what would you do? The wrong answer in this case is to ask someone in the office right away. Even though I pride myself to being a good resource for my co-workers, many times, the answers to all problems could be as simple as a google search. There should be an amount of time spent in research before they ask others to take the time to supply help. The previous question is a good lead into this because, if a developer has a good set of online resources, then they can solve most of their problems fairly easily.
  • How many web vulnerabilities are you aware of? And what can you do to prevent them? There are plenty to choose from, such as Reflected Cross Site Scripting, Persistent Cross Site Scripting, SQL Injection, Code Injection, Invalid Session management, Using default configurations, weak  encryption, improper error handling, poor input validation, weak authorization and authentication, and cross site request forgery, to name a few. The solutions to these problems include education, proper design, code review, testing, and common sense. A Good experienced programmer would recognize most, or more, of these. A mid-level programmer should know at least 4 or 5. A low level programmer would probably only recognize SQL Injection and cross site scripting.
  • Follow up question, what is the difference between Reflected Cross Site Scripting and Persistent Cross Site Scripting? Reflected is a flaw in which the user's input is reflected back to the user, this can cause a problem by allowing JavaScript code to be run in the browser. Persistent is when a similar flaw is saved to some persistent storage. Being able to answer this shows that they are not just reading buzz words.
  • Are you familiar with JavaScript? If so, how do you debug JavaScript? There are many tools out on the internet to help with JavaScript and CSS debugging. For example, I use FireBug, a plug-in for the Firefox browser to determine how JavaScript is running.

Advanced Programming:

  • Explain the benefits of Object Oriented Programming, and describe some techniques that can be used with of OOP. The benefits include improved maintainability, design, portability, modularity, and extensibility. Some techniques include data abstraction, encapsulation, modularity, polymorphism, and inheritance. As a follow up, ask to explain examples of each.
  • What are the differences between using stored procedures in a database vs writing the logic in a language like Java, and then calling the data with normal SQL queries. The DB stored procedures are compiled on the DB, so you can get better performance, but there is added complexity, since you may have different code based on different database’s.  This question shows that the candidate knows databases, and not just how to run SQL queries.
  • How would you debug a performance problem?  For example, if you have a web page that takes a long time to load, how would you debug it? This is a pretty open question. The candidate should be wondering, is this a static or dynamic page? Is there a lot of data? Large Images? Media? Try to determine where the performance problem is. Is it the Database? Check the SQL queries, analyze the database connections when the page loads. Is it the Code? Review the code for intensive loops or data calculations, watch the resources on the server, when the page loads. Is it the Network? Run traceroute or a network sniffer to look at how the packets are being transferred. All of these are potential areas for a performance problem.
  • Can you explain the differences between SOAP based web services and RESTful web services? SOAP is a standard, created by the W3C, that defines a request and response message for web services. REST is an architectural design which relies on XML, and HTTP request types. A follow up question would be to ask where you would use each.
  • Are you familiar with Aspect Oriented Programming (AOP)? If so, what benefit does it provide? Most experience developers should at least understand the concept of AOP. AOP allows for separating redundant code from modules so the developer can focus on the task at hand of the module. This saves development time, and improves maintainability and modularity.

Mid-level Questions:

  • What are 4 different types of scope used in web based programming languages like Java? Application, request, session, and page.
  • How would you loop through the GET and/or POST request parameters, when you don't know the names of the request parameters? This can be different depending on the server side language the candidate is being interviewed about, but they should all recognize the name-value pair characteristic of the request. In Java they can easily use request.getParameterNames() for an array of names. A followup question is to ask what sort of information is passed in the request. The candidate should be able to talk about the header information.
  • How would you design a multi-tiered application to separate display from business logic? This can be answered in many ways, but using a framework, an object oriented design, and or a design pattern,  can accomplish this. The key focus it to try to try and keep the HTML from getting cluttered up with logic code.
  • What software frameworks are you familiar with? Each language has their own frameworks. In Java there are quite a few, including Struts, Spring, Seam, and JavaServer Faces. A mid level programmer should be able to understand frameworks, at least on a high level.
  • What are some of the differences between Oracle and SQL Server, in terms of how the SQL is written? Most of the differences are with the out of the box functions, like getting time stamps getDate() for SQL Server, and SYSDATE for Oracle. But the SQL queries themselves are slightly different, especially with joining tables. SQL Server uses LEFT and RIGHT OUTER/INNER Join Syntax and Oracle has a short hand notation of using (+) to represent an outer join.
  • What are the benefits and disadvantages of Normalized vs. Non-normalized database design? Normalized database tables are easier to use, understand, and cuts back on redundant data. However, non-normalized table perform better.  Typical web applications use a normalized design, but reporting applications sometimes uses a non-normalized design to be able to run queries with a lot of data)

Low level programmer:

If the interviewee only answer a few of the questions above, then chances are, they are a low level programmer. If they are claiming more then 5 years experience, and cannot answer any of the above questions, that should be a red flag. However, someone with 2-4 years experience is still learning, and you could ask them how they keep up on the new technologies.

Overall, different organizations have different requirements for who they hire. Some have a strict budget to follow, others have a reputation to uphold. Regardless, the interviewer should have the knowledge of who they want on their team. If the person performing the interview cannot answer the questions they are asking, then it can be very easy to hire the wrong person.

Friday, October 22, 2010

Software Engineer Job Interviews

About 2 months ago, I was notified by a friend, about some job openings for Software Engineers, at her company.  I am pretty comfortable in my current position, but this other company looks pretty cool. Things like flex time, no dress code, high tech environment, and free coffee, makes me think that it could be the perfect job.

After a couple hours of research, it also looks like this company has high standards and employs really talented people. Not just a “Web Development” company, they expect you to understand Big O notation, and different design patterns. I consider myself an “Average Web Geek”, but I did receive a Masters in Computer Science, and my interests have always leaned more towards the engineering of software, compared to the assembly line culture that some IT companies represent.

I talked to the recruiter of this company, pretty standard stuff. She tells me that there are 4 phases to the interview process. There is first contact, that is that talk with the recruiting officer. Then there is a “homework” assignment, so the engineers can evaluate my initial coding skills. Then there is a phone interview with an engineer. Finally, there is a face-to-face. With the exception of the “homework” it is pretty standard.

Now, I should explain, I enjoy my job. But I am frustrated by the lack of a technical track. My company is looking for people to build web sites, then spend the rest of their professional lives as non-technical managers. I am also realizing there is a significant amount of co-workers who do not take as much pride as I do in engineering, and innovation. I consider myself the “Average” web geek, because I don’t see myself as anything special. Though, when I look around my office, I see a lot of people who are comfortable with their position, and the fact that they have no need to advance their knowledge, or experience. So, as happy as I am, I need to expand, and continue to advance my technical experience. This other company looks to have the direction I want to focus on.

The recruiter sends me the assignment. There are a couple questions about object oriented programming, and a Java interface for which I am required to supply an implementation. Question 1 is “What is an abstract class? Describe a common pattern that requires the use of abstract classes?” This is interesting. Normally I would use this as a question in a phone interview. It is a pretty basic question to test someone's basic object oriented knowledge. The other questions are specific to Java.

Based on the Job description, I know that they want an engineer with a large amount of Java experience, along with Spring, Hibernate, and JUnit. So, even though it isn’t required, I make sure that the code I send with the implementation class, contains these these concepts.

After a couple weeks, I hear back from them. They want to to set up a phone interview with a Lead Engineer. At this point, I am pretty excited. They seem like a pretty high tech company, with some really smart employees. So, I agree with excitement. The call comes in a couple days later. I take the call in the parking lot of a Trader Joes. I was expecting the interviewer to ask about my Java experience, shoot me some questions about Spring and Hibernate. Maybe test my knowledge of Java core concepts. But no, he hits me with a question to evaluate how I think.

He says “Lets say you are designing a card game, and you need a method to shuffle the cards. Describe the algorithm to do it.”

This question is very familiar, since I had it as a test question in my algorithms class during my graduate studies.  Unfortunately, that was about 8 years ago.  Though it shouldn’t be too hard. So, my first answer is to use 2 arrays. 1 contains 52 Card objects, and the other is empty. Then, using a random number generator, I would grab cards at random, and store them in the other array. Now, my current job hardly requires me to consider algorithms which are really speedy. Our datasets are small, and the schedules are short. So, usually the first solution is fine.  So, I was expecting his next question… “What about the blank cells in the original array? Are you concerned about those?”

Yes, of course. So, without thinking about redesigning the algorithm, I tell him that each time we remove a Card object, we can recreate the array with the size decremented. As I am talking I am realizing how stupid, and inefficient this sounds.  So I say, “It will work, but it isn’t efficient.” He then asks me to describe the performance in Big O notation. So…. there are n Card objects in the array (n = 52) , looping through the array n times to remove the Card objects. Then I would recreate the array as n-1, for each card. That is n*(n-1). Making the performance O(n^2), aka quadratic. Not very good considering that there is a better way.

After a couple minutes, it comes to me. There is actually a much simpler solution. I don’t need 2 arrays. Using just the original array of Card objects, I can generate 2 random numbers, and swap the cards. The random number generator, and swap function, would run in constant time, O(1). That leave the number of iterations. Therefore, the performance is O(n), aka linear. Much better.

The following code represents this algorithm:

public static void shuffleCards(Card[] deck) {
  Random generator = new Random();
   // shuffle the cards
   for (int i = 0 ; i < deck.length ; i ++) {
       int x1 = generator.nextInt(n) + 1;
       int x2 = generator.nextInt(n) + 1;
       Card card1 = deck[x1];
       Card card2 = deck[x2];
       deck[x1] = card2;
       deck[x2] = card1;
   }
}

This was the answer he was looking for. At this point, I could have pointed out that the code could be simplified by using the java.util.Collections class, which has a shuffle(List<?> list) method, and runs in linear time.

The phone interview continues with me asking questions about companies, and him asking a couple more questions about my experience. I guess I made an impression, since the recruiter schedules a series of face-to-face interviews. That will happen in the next week. So we would see what happens.

The moral of the story is, that there is a difference between being a average web coder, and a software engineer. I want to see what else is out there, so I have to break out of my IT shell and focus on the foundations which I learned in those Computer Science courses. It is very easy to grow comfortable and content in a stable position. But time keeps moving, and if you don’t keep learning, you can find yourself stuck. I hope one day I’ll consider myself to be more then just an Average Web Geek.

Friday, May 23, 2008

AD Authentication and Java

Well, it was only a matter of time before my job would require our Java apps to authenticate against Active Directory. For those who don't know what Active Directory (AD) is (myself included up till last year), AD is a Microsoft Windows implementation of LDAP. It is typically used in Windows 2000 based networks to tie together the standard IT resources like mail, calendars, and desktop computers. Those of us who are on a network that uses AD typically have a desktop that authenticates against AD, as well as MS Outlook for email. The good thing about a central authentication source for network credentials is that is allows for a single username and password to be used for things. Which can eventually lead to "Single Sign-On" (which will be discussed later, when I learn how to do it).

Anyway, a central source for authentication credentials also helps with web applications on that network, because now we no longer need to worry about password management, or user registration. That can reduce the project schedule by a week or two, depending on how strict your organization's password and user registration requirements are.

So, how do we do it? Well, it is actually pretty simple. Like I said, AD is basically another version of LDAP, so in Java you can use the Java Naming and Directory Interface (JNDI). There are a bunch of LDAP classes in the javax.naming.ldap package that can help. And because the Java API is so robust, it gives you a ton of flexibility to customize your code as much as possible. At the same time, it can seem a bit intimidating. Sun has some pretty good information on their website about LDAP authentication, which can also be used for AD Authentication. Lets take a look at some code.
Hashtable env = null;
DirContext ctx = null;
boolean isAuthenticated = false;

try {
try {
String loginId = "yourdomain\\avgwebgeek";
env = new Hashtable();   // hash table for your LDAP properties

// set up the LDAP properties
env.put(Context.INITIAL_CONTEXT_FACTORY, "com.sun.jndi.ldap.LdapCtxFactory");
env.put(Context.PROVIDER_URL, "ldap://myadserver.mydomain.com:389");

// Set the authentication mechanism to be simple
env.put(Context.SECURITY_AUTHENTICATION, "simple");

// login credentials
env.put(Context.SECURITY_PRINCIPAL, loginId);
env.put(Context.SECURITY_CREDENTIALS, password);

// the following is helpful in debugging errors from the
// AD server side of things
//env.put("com.sun.jndi.ldap.trace.ber", System.err);

// Create the initial directory context
ctx = new InitialDirContext(env);
isAuthenticated = true;
} catch (AuthenticationException e) {
// this exception is typically found with incorrect credentials
e.printStackTrace();
} catch (NamingException e) {
e.printStackTrace();
}
} catch (Exception e) {
e.printStackTrace();
} finally {  
try {
ctx.close();
}catch (NamingException e) {
e.printStackTrace();
}
}

if (isAuthenticated) {
System.out.println("I am authenticated.");
} else {
System.out.println("I am not authenticated.");
}

This is a very simple application. It does the job, however, it is not very secure. The problem is that I am setting the Context.SECURITY_AUTHENTICATION to "simple". What this means is that your password is being sent over the network through clear text. If someone is running a network sniffer, they can read your password from your IP packets. One way around this is to make sure that your LDAP connection is handled over SSL. In other words, use "ldaps://" with port 636, and not "ldap://" on port 389.

If your server is not setup to handle SSL, or you just want a little extra security, you can change the Context.SECURITY_AUTHENTICATION to "DIGEST-MD5". However, as I found out, AD treats this a little differently. The concept of using "yourdomain\\yourusername" needs to be reduced to just "yourusername". It has something to do with how AD 2003 sets up the HASH for the MD5. This method will use a Hashing algorithm to be verified with the server, that you indeed know your password.

So, that's it for AD authentication. Like I said, the Java API can let you do a lot more. You can set up controls to do searches, find specific user information, and find groups assigned to users. AD groups can be used to help with the Authorization portion of your apps. Take a look at the javax.naming.ldap API, there is a lot there, but it isn't too hard to follow.

Monday, April 28, 2008

PHP and LDAP

The joys of authentication. So, now I am working on running some of our PHP apps to authenticate against Active Directory (AD). Well, my limited knowledge in AD is telling me that it shouldn't be too hard. After all, AD is just a Microsoft interface to LDAP. Right?

Well, it turns out that, yeah, from a PHP perspective, it is. It is actually pretty easy, once you are comfortable with working with an LDAP structure.

So, here is the basics of it. I have a simple page that all it does is print out the results of a simple authentication to an AD server:
<html>
<body> 
<h1>AD Test</h1>
<?php
// Variables to use with ldap_bind
$ldapuser  = 'username@some.domain.com';     // ldap username with suffix
$ldappasswd = 'notmyrealpw';  // associated password

// connect to AD server
$adconn = ldap_connect("ad_controler.myco.com")
or die("Could not connect to LDAP server.");

// if a connection was made attempt a binding
if ($adconn) {

// bind to ldap
$ldapbind = ldap_bind($ldapconn, $ldapuser , $ldappass);  

// Check authentication
if ($ldapbind) {
echo "User is authenticated.";
} else {
echo "User was not authenticated.";
}

// unbind the connection
ldap_unbind($adconn);
}
?>
</body>
</html>


Basically, you have your username, for example avgwebgeek. Then you have a suffix, which is your account suffix for your domain, for example "@myco.com". Put them together (along with a password) and you have your AD authentication information. Of course, the username and password would be sent through request parameters. Granted, this is a pretty simplistic view of AD authentication. You can do a whole lot more, like searching, and updating AD information. I learned all about it by looking through the adLDAP Project. The hard part was realizing that I didn't have PHP enabled on my server. Hint: if you get a message like "Call to undefined function: ldap_connect() ", that means that your LDAP isn't enabled for PHP. The adLDAP page has a FAQ that explains that as well.

Wednesday, April 23, 2008

Where’s my data coming from?????

I work on applications that interface with other applications using Web Services, and XML feeds. These apps help track employee information for different companies to track thing like pay, performance, and security issues. The problem is, I don’t know where the actual content is originating from. That can be problematic, because you are assuming that your data provider will let you know if there is a format change or discontinuation coming down the pipe. Typically, the provider will let you know, somehow. However, in the case of systems “piggy backing” on data of other systems, then in turn, providing that data to further systems, eventually, the message of change gets lost.

I found this out the hard way a couple weeks ago. Luckily for me, it was for a low priority system. I manage a series of blogs for a corporation. I don’t manage the content, but I taker care of the themes and plugins. One of the plugins that I wrote was a “Company Updates” plugin, basically an RSS type of news reader that uses plain XML, and not the RSS standard. The data was coming from what I thought to be the HR department. However, they were just sending it to me through a XML feed, as they were getting it from a small group in the company who were generating the data by hand. Well this group changed the format of the XML tags, and notified HR. But HR didn’t notify me. All of a sudden, I come to work and all of the Blog users were calling me because the “Company Updates” feed was busted. I noticed the change in the XML feed, so I fixed the problem. However I couldn’t understand why I wasn’t notified.

I contacted HR, and they didn’t know anything because they never wrote the Change notification down. Apparently the employee who was notified about the change assumed that in my super natural power over XML feeds, my plugin would have updated dynamically with the XML change. True, if I had a DTD or Schema, I could have written a more intelligent plugin, but I didn’t have a DTD or Schema, and this was a pretty small system. In this case, it is easier to modify the code as changes occur. Anyway, the HR people got me in touch with the group who produces the feed. They told me about the notifications. Needless to say, I switched the origin of the feed to this group.

The moral of the story is that documentation is a REALLY good thing. In this case the problem was small, and easily rectified. However, it made me review my other interfaces, and I decided to write up a Provider/Consumer agreement about data, so the actual origin of the data is known. This agreement is review by all parties in involved and helps to keep me in the loop when changes happen.

Wednesday, January 16, 2008

Sun buys MySQL

I just read on Slashdot that Sun Microsystems acquired MySQL. That is of particular interest to me because I love MySQL. I use it for all of my side projects and some of my professional projects. It should be interesting because MySQL has gained a lot of popularity with PHP developers and the whole LAMP (Linux, Apache, MySQL, and either Perl, Python, or PHP) movement. Many web developers who use MySQL, don't seem to think of Java as the first language to use with it. Java, being Sun's major programming platform, is now going to be linked to MySQL. So I wonder if the LAMP acronym might be changed to LAMJ???? I know, that doesn't make much sense. But from my perspective, a web developer who loves to develop in Java, that combination opens a lot of doors.

This acquisition might also open MySQL up to a more professional world. For the past 10 years, though MySQL has been used in some minor projects and applications, the big question I always see in regards to databases is "Oracle or SQL Server?" Many of the organizations I worked for didn't really trust MySQL. So now, I wonder if companies and organizations might change their mind, now that MySQL has Sun behind them.

Tuesday, December 18, 2007

What happened to javarss?

About a year ago, javarss.com got hacked. Back up, for those who do not know that Javarss is, it was a simple site that had a really nice AJAX driven interface containing RSS feeds from all of the major Java websites. It was simple, just a bunch of links, organized by Website, and when you hovered over them, you'd get a brief synopsis of what the article was about. A one stop shop for Java info and advancements. For busy people, like me, it saved me the trouble of surfing through a dozen or so web sites each day to to find interesting topics.

Well, it was great, until it got hacked back in March 07. The hacker go into the server, and replaced the javarss page with a evil (and kind of neat) looking page, proclaiming "Hacked By Secretary Lucifer". The hacker even left a couple email where the server admins can contact them, possibly about getting their access back. Well, a couple months later, the server admins finaly gain control again, but they never bring Javarss back up. All that remains is a message that reads "We are restoring the site. Please visit us later...". This has been like this for about 6 months. I would love to know how the hackers got in, I love learning about that stuff.

So, what happened? I mean, c'mon, they are just loosing money, and loosing loyal followers with each passing day that they stay closed. I can honestly say that since their demise, I have started to frequent DZone a lot more. DZone is a similar concept, but with a much more rich interface, and they are find articles for languages other then Java. After all, we are developers, not Java Developers. So, I highly recommend you check it out if you haven't already.

Friday, November 16, 2007

Uggg, blogging is hard

Wow, I haven't posted since January. I actually forgot about this blog for a while. I was about to delete it, figuring that no one is reading it. Besides, I never post to it. Is it because I don't have enough time? Nah, I don't think that is the problem. I think that it has a lot to do with the fact that I wanted this blog to be something epic. I wanted to reach the average coders out there who are too timid to risk being teared appart on a Linux user forum by asking a question about PHP. But I wanted it to be informative to everyone like Ben Forta's blog, or the Coding Horror blog (seriously, you should check this one out). And in doing so, I was unconsciously putting the concept of this blog up on a throne. The importance I placed in my message is what was causing me to delay my next entry. I was too nervous that what I would write wouldn't be important enough, or would be considered as child's play to some of the more advanced coders out there.

But I've realized something over the past couple months. Given my experience and education, I am actually on the same playing ground as many of the advanced coders. They just have more confidence in what they say than I do. So, the trick is to post blogs, and keep posting. Not worrying about whether or not my information is going to be an epic, or an awe inspiring message of programming genius. Hell, if I can help one person set up Tomcat with a project that is a non-war deployment, and have their servlets and JSPs reloadable, then I think I've done my job.

That reminds me, along with my previous posts about Tomcat 5.5, make sure you set the reloadable attribute to true in the context.xml file under the $CATALINA_HOME/conf directory.

<context reloadable="true">

That will aid in making your classes (servlets) reloadable without restarting the server, or using WAR file deployment.

Till the next time. Keep coding.

Friday, January 05, 2007

Dead Winter Dead and Web Application Security

Wow, I haven't added to this blog in over a month. The ironic thing is that the winter time is usually the slowest time for a developer who works for the government. Especially where I work. The organization that I work for are re-organizing themselves, and have no time to look at proposals or statements of work. So, we just sit and wait for something exciting to happen. Since we are on a fixed price contract, we still get paid.

Although, times haven't been that slow for me at work, and home. My wife, kids, and myself have all been sick. The relatives came in from out of state. And the holiday season in general makes simple tasks a mad house. But enough of this non-technical stuff. Anything interesting?

Well, yeah, kind of. I was assigned a task to look into new security requirements for many of our apps. These requirements were signed by President Bush about 2 years ago. Some are simple, like session timeouts must be set at 30 minutes, and tracking all access to "Personally Identifiable Data" (aka Social Security Numbers). Some were a little more complex, like enforcing 2-factor authentication. Since many of our apps deal with secure data, and because of the whole VA Laptop thing that happened last summer. My group needed to do something, fast. For all new apps, the security requirements have already been put in place, but the older apps needed to be updated.

The interesting part was working on a solution for enforcing 2-factor authentication. Typically, a website has 1-factor authentication, a password to proove that the user is valid. 2-factor authentication is a way of not only proving the user is valid with something that they know, but to request something that they have as well. Things like smart cards, biometrics, and security tokens are just a couple ways to help prove that you are who you say you are. Our government client decided on using security tokens to solve this problem, in the short term, at least.

The security token works on the concept of this little device that displays a 6 digit number, that changed every minute. A user would enter a pin and the 6 digit pass code to prove their identity. However, these little devices are about $70 a pop. Considering that hundreds to thousands of people use our systems, that can get costly. Well, it turns out that most people who use the systems, don't need to have access to the secure data anyway, they only need access to the non-secure data on the sites. We have user roles to enforce who can see what. However, there can be 2 people on the same role who can see the same data, but only one needs to see the secure data. We wanted a simple way of hiding the secure data, without recreating the users into new roles, and causing as little modification to the code as possible. Therefore, only a small fraction of people will need the tokens.

All of our security and user authentication code was generated through ColdFusion/Java and an Oracle back end. The use of the ISAPI dll that came with the secure tokens wasn't useful for us, and I couldn't find anything on the net about this type of implementation. So I had to design our own way of doing this. I will discuss my solution in my next post. If anyone has run though this situation before, I would love to hear how you handled it.

Thursday, November 30, 2006

Pretty Good Combobox HTML Implementation

I saw this a couple months ago and love it. It is a drop down that you can enter text directly in. Pretty clever, they used CSS to put a textbox on top of a drop down. Then using a little JavaScript, the two become one.

Here's the link: http://particletree.com/features/upgrade-your-select-element-to-a-combo-box/

Tuesday, November 07, 2006

News About Programmer Jobs

I read in the October 16, 2006th copy of Information Week, that jobs for coders is starting to stabilize. The article has a dark view on this news, mainly because the amount of jobs is not going up. However, I see it differently. Overall, since 2000, the number of employed programmers has dropped about 200,000. While the number of people who considered themselves programmers, has dropped as well. I started my professional career in 1998 as a Java developer for an IT consulting firm. We got torn to shreds when the dot com crash of 2000/2001 hit. I was one of the lucky ones who managed to stay employed. I saw a lot of friends out of work. Every year it seemed like things were getting worse.

So, now, with the programming jobs becoming more stable, things look much better. I would rather that the jobs remain stable, then go down hill. The article lists updating legacy systems, customizing apps, and writing new apps as positions where developers are in high demand. One thing is certain though, technological progress will continue, therefore, programmers, and software engineers, will always be needed.

Friday, November 03, 2006

Reading Files With ColdFusion, or Java???

One project I was on, not to long ago, consisted of transferring a datafile between 2 systems. Unfortunately, this datafile wasn't XML. It was a pipped delimited file of product records. The main purpose of this file was to sync up data between the two systems. Therefore, I needed to read the file, line by line, parse out each field of the file, and compare it with current records in the database. Since the language of choice for this particular project was ColdFusion, I was encouraged to keep with the standard.

Being a Java guy, with some php and perl background, I knew I could adapt. But how can I make this as efficient as possible. The app server we were using was ColdFusion version 6.1. So, I know I had options. At least it wasn't version 5 or below, like some other organizations who shall remain nameless.

Anyway, ColdFusion has a tag called <cffile>. This tag is nice and easy to use. However, it has one main drawback to it. It holds the entire file in memory at once. It might not seem like a problem, but I am talking about somewhere around 200,000 records. With each record containing 80 fields. That can make the application more prone to memory errors. Especially, if this is a user requested job, and not a scheduled job.

However, the standard Java development kit has the FileReader, used with a BufferedReader, that can stream the file, line by line. So, I figured I'd give it a try. I started by comparing the two methods of file reading. I wrote a sample ColdFusion page which would read the file in and compare the time it takes to read/process the file using the <CFFile>, and the Java FileReader methods. I started with a small file, containing 200 records. Then moved up to a file with 25,000 records. The code is as follows:
Using cffile:
<cfset filePath = "./DataFile200Records.txt">
<cffile action="read" file="#filePath#" variable="fileContents">

Using CFFile<br>
<cfoutput>
Start:#now()#<br>
<cfloop index="line" list="#fileContents#" delimiters="#Chr(10)#">
<!--- Account for empty values --->
<cfset newline = " " & Replace(line, "||", " | | ", "all") & " ">
<!--- Convert the pipped delimited list to an Array --->
<cfset fieldArray = listtoarray(newline,"|")>
<cfloop index="i" from="1" to="#arraylen(fieldArray)#">
<!--- Loop through each field --->
<!-- #fieldArray[i]# *** -->
</cfloop>
</cfloop>
End:#now()#<br> <!--- End Timestamp --->
</cfoutput>

Using Java FileReader:
Start:<cfoutput >#now()#</cfoutput><br><!--- Start Timestamp --->
<cfscript>
// create a FileReader and BufferedReader objects
fileReader = createObject("java","java.io.FileReader");
buffReader = createObject("java","java.io.BufferedReader");

// Instantiate the FileReader with the file path
fileReader.init(filePath);

// Instantiate the BufferedReader with the fileReader Object
buffReader.init(fileReader);

// read the first line into a String
fileLine = buffReader.readLine();

// Loop while the String is defined
while (isDefined("fileLine")) {
// Account for empty values
newline = " " & Replace(fileLine, "||", " | | ", "all") &amp;amp;amp;amp; " ";
// Convert the pipped delimited list to an Array
fieldArray = listtoarray(newline,"|");
for (i = 1; i lt arraylen(fieldArray); i = i+1)
{
// Loop through each field
writeOutput("<!--" & fieldArray[i] & " *** -->");
}
// read the next line to continue the loop
fileLine = buffReader.readLine();
}

// close the Reader objects
buffReader.close();
fileReader.close();
</cfscript>
End:<cfoutput >#now()#</cfoutput><!--- End Timestamp --->
Essentially, each section of code is doing the same thing. It reads the file, line by line. Parses the fields, including empty fields. And it loops through each field, displaying it as an HTML comment. The main difference is that the Java FileReader and BufferedReader objects have to be closed when they are no longer needed. You always want to close a Streamed object, otherwise it can take up memory.

So, what are the results? You would think that CFFile would run faster, being that the file is in memory. And memory access is usually pretty quick. Well, for the small, 200 record, file, that is correct. Below are the results:
200 RecordsUsing CFFile:
Start:{ts '2006-11-03 10:23:32'}
End:{ts '2006-11-03 10:23:33'}

Using Java FileReader
Start:{ts '2006-11-03 10:23:33'}
End:{ts '2006-11-03 10:23:35'}

Using the Java objects took a second longer. I think that may have to do with the overhead involved with opening and closing the Reader objects. But that would be fairly the same delay on other files, regardless of the size of the file.

Now, for the file with 25,000 records:
25,000 recordsUsing CFFIle:
Start:{ts '2006-11-03 10:47:46'}
End:{ts '2006-11-03 10:47:55'}

Using Java FileReader
Start:{ts '2006-11-03 10:47:55'}
End:{ts '2006-11-03 10:47:59'}

After 25,000 records, the CFFile took 9 seconds, and the Reader objects took 4 seconds. That is quite a difference. So, we ended up using the Java code, and it works well. We did some other mods, like using the <cftry> and <cfcatch> blocks to handle errors gracefully. If this was a full Java approach, I would also use the finally {} block to close the Reader objects. The moral of the story is, we have options.

Tuesday, October 31, 2006

More about Tomcat Contexts

Here is some more relevent information about Contexts in Tomcat. This is just an addition to the post I made yesterday. I found this on the Tomcat Deployer How-To webpage.

A Word About ContextsIn talking about deployment of web applications, the concept of a Context is required to be understood. A Context is what Tomcat calls a web application.

In order to configure a Context within Tomcat a Context Descriptor is required. A Context Descriptor is simply an XML file that contains Tomcat related configuration for a Context, e.g naming resources or session manager configuration. In earlier versions of Tomcat the content of a Context Descriptor configuration was often stored within Tomcat's primary configuration file server.xml but this is now discouraged (although it currently still works).

Context Descriptors not only help Tomcat to know how to configure Contexts but other tools such as the Tomcat Manager and TDC often use these Context Descriptors to perform their roles properly.

The locations for Context Descriptors are;
  1. $CATALINA_HOME/conf/[enginename]/[hostname]/context.xml
  2. $CATALINA_HOME/webapps/[webappname]/META-INF/context.xml
If a Context Descriptor is not provided for a Context, Tomcat automatically creates one and places it in (1) with a filename of [webappname].xml although if manually created, the filename need not match the web application name as Tomcat is concerned only with the Context configuration contained within the Context Descriptor file(s).

Monday, October 30, 2006

Tomcat JSP and Servlet Reloading

Well, for my first official blog post, I'd like to talk about a problem I've recently ran into with JSP loading in Tomcat 5.5. Now, I've used Tomcat for years, but never took the time to learn it totally inside and out. I've become very good at using Tomcat over the years by tackling problems as they arise. So, I am very comfortable with it.

So, last week, when I uploaded a JSP to my Tomcat web application directory, and it didn't show the changes I've made to it, I was perplexed. This was the most cut and dry part of using Tomcat. Uploading JSPs and having them dynamically reload. Why was this happening? Why? Why? Why????? After researching the problem, I found that others had the same problem, but their questions were never answered, or the solutions were not available on the web. So hopefully, this post will help someone.

First, I will describe the environment. This is a server that was created for students, at the university that I teach at, to use for assignments. It is an Apache Tomcat version 5.5.17 application server, on a Sun Sparc OS 5.10. Since this is an intro to dynamic web development class, the students do not deploy WAR files. Rather, they deploy their servlets and JSPs to a pre-created web application directory, under the tomcat "webapps" directory. When I created this directory, I did not change any Tomcat configuration files. I just created the root directory "classfiles", and the WEB-INF, classes, and lib directories.

The first problem was loading servlets. That one was an easy one. The web.xml file needs to tell Tomcat what to look for and how to load servlets. So, I made a basic web.xml file under the directory "classfiles/WEB-INF" that included the following code in it:
<?xml version="1.0" encoding="ISO-8859-1"?>
<!DOCTYPE web-app
PUBLIC "-//Sun Microsystems, Inc.//DTD Web Application 2.3//EN"
"http://java.sun.com/dtd/web-app_2_3.dtd">
<web-app>
<servlet>
<servlet-name>invoker</servlet-name>
<servlet-class>
org.apache.catalina.servlets.InvokerServlet
</servlet-class>
<init-param>
<param-name>debug</param-name>
<param-value>0</param-value>
</init-param>
<load-on-startup>1</load-on-startup>
</servlet>
<servlet-mapping>
<servlet-name>invoker</servlet-name>
<url-pattern>/servlet/*</url-pattern>
</servlet-mapping>
</web-app>
This tells tomcat to look for anything that has the /servlet/ in the user for the "classfiles" application and load it as a Servlet. For example: http://localhost:8080/classfiles/servlet/HelloWorldServlet

Easy enough, however, the dynamic reloading of the servlets wasn't happening. After looking through the documentation on the Tomcat site, I found out that Tomcat has to be explicitly notified about web applications, and if they should be reloadable. By default, if I was to use WAR file deployment, Tomcat would handle this automatically. Since this is not the case, I need to explicitly tell Tomcat what to do. This can be done either by modifying the server.xml file, in the tomcat "conf" directory. Or, to give me more direct control over the web application, and reduce the risk of accidentally screwing up someone else's context, I decided to make my own context configuration. This is done by creating a new directory under your web application directory called "META-INF". Then creating a file in there called "context.xml". Therefore, the contents of my file at "classfiles/META-INF/context.xml" was as follows:
<?xml version='1.0' encoding='utf-8'?>
<Context displayName="Software Engineering Student Projects"
reloadable="true">
</Context>
I am telling Tomcat to monitor the WEB-INF/classes and WEB-INF/lib directories for any changes. And WHAM, the servlets, packages, and other classes are now reloadable. But what about JSPs??? Am I missing something? Aren't JSPs just simple servlets? Well, yeah, sort of. It turns out that you need to set the docBase for the JSPs to reload. So now my context.xml file looks like this:
<?xml version='1.0' encoding='utf-8'?>
<Context displayName="Software Engineering Student Projects"
docBase="/usr/local/tomcat/classfiles" reloadable="true">
</Context>
And now my JSPs are reloading with no problems.

So the context.xml file can be used to make your pages, servlets, classes, and packages reloadable, if you are not using the default deployment capabilities of Tomcat. Tomcat does all this stuff automatically when deploying applications as a Web ARchive, WAR file. Is that it??

"But that is not all I can do" said the cat. Through research into the context.xml, and server.xml, files, you will find a ton of more capabilities that these files provide. Such as data sources, cross context sessions, and application level variables. But that is a post for another time.

For more information, check out the Tomcat documentation on configuring the Context container.