Showing posts with label mongo. Show all posts
Showing posts with label mongo. Show all posts

Friday, March 22, 2013

Installing (but not configuring) the broker service by hand

I'm working through a totally(?) manual installation of the OpenShift Origin service on Fedora 18. The last post on this topic was about building the RPMs on your own Yum repository. This time I'm going to install the broker service and make a few tweaks that are still required.

One seriously major thing to note is that I don't recommend actually doing this. I'm doing it to shed some light on some of the things still going on in the development process and to highlight the ways in which you can get some visibility into the installation and monitoring of the service.

If you're interested in building and running your own development environment or service for real, I suggest starting by reading through Krishna Raman's article on creating a development environment using Vagrant and Puppet and the puppet script sources themselves to see what's involved.  Finally there's a comprehensive document that describes the procedure with fewer warts.


Ingredients

As usual, I start with a clean minimal install of Fedora 18.  In addition this time I also have a yum repository filled with a bleeding-edge build from source as I described previously.  Finally I have a prepared MongoDB server waiting for a connection.

I'm replacing my real URLs and access information with dummies for demonstration purposes.


  • Yum repo URL
    http://myrepo.example.com/origin-server
  • MONGO_HOST_PORT="mydbhost.example.com:27017"
  • MONGO_USER="openshift"
  • MONGO_PASSWORD="dontuseme"
  • MONGO_DB="openshift"

Preparation

Since I'm building my own packages from source and placing them in a Yum repository, I need to add that repo to the standard set. I'll add a new file to /etc/yum.repod.d referring to my yum server.

Even if you're building from your own sources, there are still some packages you need to get that aren't in either the stock Fedora repositories or in the OpenShift sources. These are generally packages with patches that are in the process of moving upstream or are in the acceptance process for Fedora. Right now a set is maintained by the OpenShift build engineers. I need to add the repo file for that too:

[origin-server]
name=OpenShift Origin Server
baseurl=http://myrepo.example.com/openshift-origin
enable=1
gpgcheck=0
[origin-extras]
name=Custom packages for OpenShift Origin Server
baseurl=https://mirror.openshift.com/pub/openshift-origin/fedora-18/x86_64/
enable=1
gpgcheck=0
At this point you can install the openshift-origin-broker package.
yum install openshift-origin-broker
...
  urw-fonts.noarch 0:2.4-14.fc18                                                
  v8.x86_64 1:3.13.7.5-1.fc18                                                   
  xorg-x11-font-utils.x86_64 1:7.5-10.fc18                                      

Complete!


There are a set of Rubygems that are not yet packaged as RPMs. I need to install these as gems for now.

gem install mongoid
Fetching: i18n-0.6.1.gem (100%)
Fetching: moped-1.4.4.gem (100%)
Fetching: origin-1.0.11.gem (100%)
Fetching: mongoid-3.1.2.gem (100%)
Successfully installed i18n-0.6.1
Successfully installed moped-1.4.4
Successfully installed origin-1.0.11
Successfully installed mongoid-3.1.2
3 gems installed
Installing ri documentation for moped-1.4.4...
Building YARD (yri) index for moped-1.4.4...
Installing ri documentation for origin-1.0.11...
Building YARD (yri) index for origin-1.0.11...
Installing ri documentation for mongoid-3.1.2...
Building YARD (yri) index for mongoid-3.1.2...
Installing RDoc documentation for moped-1.4.4...
Installing RDoc documentation for origin-1.0.11...
Installing RDoc documentation for mongoid-3.1.2...
There are a number of gem version restrictions in the broker Gemfile which are not met by the current rubygem RPMs.  I have to remove the version restrictions so that the broker application will use what is available. This risks breaking things due to interface changes, but will at least allow the broker application to start.

sed -i -f - <<EOF /var/www/openshift/broker/Gemfile
/parseconfig/s/,.*//
/minitest/s/,.*//
/rest-client/s/,.*//
/mocha/s/,.*//
/rake/s/,.*//
EOF


For some reason, even with the --without clause for :test and :development, bundle still wants the mocha rubygem.  This should not be required for production, but right now you need to install it so that the Rails application will start.

yum install rubygem-mocha
...
Installed:
 rubygem-mocha.noarch 0:0.12.1-1.fc18

Dependency Installed:
  rubygem-metaclass.noarch 0:0.0.1-6.fc18

 Verifying The Dependencies

Now that all of the software dependencies have been installed (mostly by RPM requirements through Yum, and finally through gem requirements and some version tweaking of the Gemfile) I can check that all of them resolve when I start the application. Rails will call bundler when the application starts so I'll call it explicitly before hand. I'm only interested in the production environment, so I'll explicitly exclude development and test.

cd /var/www/openshift/broker
bundle --local
Using rake (0.9.6) 
Using bigdecimal (1.1.0)
....
Using systemu (2.5.2)
Using xml-simple (1.1.2)
Your bundle is complete! Use `bundle show [gemname]` to see where a bundled gem is installed.

If I try to start the rails console now, though, I'll be sad. It won't connect to the database.

Configure MongoDB access/authentication

The OpenShift broker is (right now) tightly coupled to MongoDB. Recently it switched to using the rubygem-mongoid ODM module (which is a definite plus if you have to work on the code).

The last thing I need to do before I can fire up the Rails console with the broker application is to set the database connectivity parameters. One side effect of using an ODM is that it establishes a connection to the database the moment the application starts.

NOTE: when this is done I will not have a complete working broker server. I still need to configure the other external services: auth, dns and messaging.

Set the values listed in the Ingredients into /etc/openshift/broker.conf.

/etc/openshift/broker.conf
...
# Eg: MONGO_HOST_PORT="<host1:port1>,<host2:port2>..."
MONGO_HOST_PORT="mydbhost.example.com:27017"
MONGO_USER="openshift"
MONGO_PASSWORD="dontuseme"
MONGO_DB="openshift"
MONGO_SSL="false"
...

Now I can try starting the rails console. It should connect to the mongodb and offer an irb prompt:


To verify the database connectivity, take a look at this recent blog post.

Next up is configuring each plugin, one by one.

Gist Scripts

I'm trying something new.  Rather than including code snippets inline, I'm going to post them as Github Gist entries.

References

Thursday, March 21, 2013

Verifying the MongoDB DataStore with the Rails Console: Mongoid Edition

A few months ago I did several posts about how to verify the operation of the back end services of an OpenShift Origin broker service.   Today I discovered that this one (mongod) is obsolete.

The data store behind the broker is a MongoDB.  That one back end service isn't pluggable.  It's actually been made more tightly coupled to Mongo, but in this case that's a good thing.  What changed is that all of the Rails application model objects have been converted to use the Mongoid ODM rubygem.  All of the object persistence is now managed in the background and all of the logic can just deal with the objects as... well... objects.

There are a couple of implications for broker service verification.

  1. The broker connects to the database on startup
    This means that if the database access/auth information is wrong, the rails app will fail to start.
  2. The only simple way to test the connection is to create an object and observe the database.
    This is both simpler to do, and potentially more difficult to diagnose on failure.

I think the second point won't be as much of a downside as I would fear at first.  I suspect that if connectivity is good, the rest will be.  If it's not, it will be fairly clear why.

Configuring the Broker Data Store


Configuring the datastore access information hasn't changed.  The configuration information is still stored in /etc/openshift/broker.conf. The settings all have the MONGO_ prefix:

MONGO_HOST_PORT="data1.example.com:27017"
MONGO_USER="openshift"
MONGO_PASSWORD="dontuseme"
MONGO_DB="openshift"
MONGO_SSL="false"

Adjust these for your mongodb implementation. Remember to open the firewall for the broker on your database host.  Configure the database to listen and test the connectivity locally.

Verifying Simple Connectivity


You also want to check the connectivity from your broker host before trying to fire up the broker itself.

broker> echo "show collections" | mongo --username openshift --password dontuseme data1.example.com:27017/openshift
MongoDB shell version: 2.2.3
connecting to: data1.example.com:27017/openshift
system.indexes<
system.users
bye

You can do this repeatedly and observe the mongodb log on the database host.

Observing the Mongo Database Logs

On the database host, take a look at the mongodb logs. You should see a new entry (successful or failed) each time a client connects.

data1> tail /var/log/mongodb/mongodb.log
Thu Mar 21 20:24:26 [conn15] authenticate db: openshift { authenticate: 1, nonce: "20d6f85f33f03dee", user: "openshift", key: "60639c7ce56851a25be56bcebd98c3ed" }

Starting the Rails Console


Now that you're sure that the database is running and accessible from your broker host you can try firing up the Rails console. This assumes that you've resolved all of the gem requirements. If not, the Rails console will complain about them and exit.

broker> cd /var/www/openshift/broker
broker> rails console
Loading production environment (Rails 3.2.8)
irb(main):001:0>

If you go this far you should have seen one more authentication log record on the mongodb server. (see above)

Create a Database Object


Now we can create a CloudUser object and watch it appear in the database

irb(main):001:0> user = CloudUser.create(login: "testuser")
=> #<CloudUser _id: 514b6f6cf3da7fa491000001, created_at: 2013-03-21 20:37:00 UTC, updated_at: 2013-03-21 20:37:00 UTC, login: "testuser", capabilities: {"subaccounts"=>false, "gear_sizes"=>["small"], "max_gears"=>100}, parent_user_id: nil, plan_id: nil, pending_plan_id: nil, pending_plan_uptime: nil, usage_account_id: nil, consumed_gears: 0>

You can see that this is more than your typical Ruby object. The ID and created_at and updated_att fields are artifacts of the ODM persistence.  You won't see another log message because the database connection is persistent using the ODM.  You will find that there's now a document in the openshift.cloud_user collection.

data> echo "db.cloud_users.find()" | mongo --username openshift --password dbsecret localhost/openshift
MongoDB shell version: 2.2.3<
connecting to: localhost/openshift
{ "_id" : ObjectId("514b6f6cf3da7fa491000001"), "consumed_gears" : 0, "login" : "testuser", "capabilities" : { "subaccounts" : false, "gear_sizes" : [ "small" ], "max_gears" : 100 }, "updated_at" : ISODate("2013-03-21T20:37:00.546Z"), "created_at" : ISODate("2013-03-21T20:37:00.546Z") }
bye


Removing the Test Object


Cleaning up is just as easy

irb(main):002:0> user.delete<
=> true

And to verify that it's been removed:

echo "db.cloud_users.find()" | mongo --username openshift --password dbsecret localhost/openshift
MongoDB shell version: 2.2.3
connecting to: localhost/openshift
bye

At this point you know both that your database is running and that the broker application can connect and read and write it.

Much simpler with an ODM.

References:


Thursday, December 6, 2012

Verifying the MongoDB DataStore with the Rails Console

UPDATE: The broker data model has switched to using the Mongoid ODM rubygem.  This significantly improves the consistency of the broker data model and simplifies coding broker objects.  It also obsoletes this post.

See the new one on Verifying the Mongod DataStore with the Rails Console: Mongoid Edition


In the last post I showed how I'd verify the configuration of the OpenShift Bind DNS plugin using the Rails console.  In this one I'll do the same thing for the DataStore  back end service. (not strictly a plugin, but hey...).

DataStore Configuration


Right now the DataStore back end service is not pluggable.  The only back end service available is MongoDB. I've posted previously on how to prepare a MongoDB service for OpenShift.  Now I'm going to work it from the other side and demonstrate that the communications are working.

Since the DataStore isn't pluggable, it isn't configured from the /etc/openshift/plugins.d directory.  Rather it has it's own section in the /etc/openshift/broker.conf (or broker-dev.conf).

This is just the relevant fragement from broker.conf

...
#Broker datastore configuration
MONGO_REPLICA_SETS=false
# Replica set example: "<host-1>:<port-1> <host-2>:<port-2> ..."
MONGO_HOST_PORT="data1.example.com:27017"
MONGO_USER="openshift"
MONGO_PASSWORD="dbsecret"
MONGO_DB="openshift"
...

These are the values that will be used when the broker application creates an OpenShift::DataStore (and OpenShift::MongoDataStore) object.

The DataStore: Abstract and Implementation

At one point the OpenShift::DataStore was intended to be pluggable.  At some point the concept of an abstracted interface was dropped and the tightly bound MongoDB interface was allowed to grow organically. The remains of the original pluggable interface are still there.  Both source files live now in the openshift-origin-controller rubygem package.


The OpenShift::DataStore class still follows the plugging conventions.  It implements the provider=() and instance() methods.  The first takes a reference to a class that "implements the datastore interface" and the second provides an instance of the implementation class all pre-configured from the configuration file.

Observing MongoDB


Unlike named MongoDB writes to its own log by default.  The logs reside in /var/log/mongodb/mongodb.log. (as controlled by the logpath setting in /etc/mongodb.conf.) Verbose logging is controlled in the mongodb.conf as well.  For this demonstration I'm going to enable that by uncommenting the line in /etc/mongodb.conf and restarting the mongod service.

MongoDB also has a command line tool that can be used to interact with the database as well.  The CLI tool is called mongo. I can invoke it like this:

mongo --username openshift --password dbsecret data1.example.com/openshift
MongoDB shell version: 2.0.2
connecting to: data1.example.com/openshift
> show collections;
system.indexes
system.users

This shows an initialized database, but no OpenShift data has been stored yet. The two existing collections are the system collections.  OpenShift will add collections as needed to store data.

With these two mechanisms I can observe and verify access and updates from the broker to the database through the OpenShift::DataStore object.

Creating an OpenShift::DataStore Object


I'm going to create the OpenShift::DataStore object in the same way I did with the OpenShift::DnsService object.  I call the instance() method on the OpenShift::DataStore object.

cd /var/www/openshift/broker
rails console
Loading production environment (Rails 3.0.13)
irb(main):001:0> store = OpenShift::DataStore.instance
=> #<OpenShift::MongoDataStore:0x7f9e42ed4918
 @host_port=["data1.example.com", 27017], @db="openshift", @user="openshift",
 @replica_set=false, @password="dbsecret",
 @collections={:application_template=>"template", :user=>"user", :district=>"district"}>
irb(main):002:0>

Now I have a variable named db which contains a reference to an OpenShift::MongDataStore object. I can see from the instance variables that it is configured for the right host, port, database, user etc.

Checking Communications: Read


Now that I have something to work with its time to see if it will talk to the database.

The interface to the DataStore is much more complex than the DnsService interface is.  Since we're only checking connectivity that's not a problem.  Once I've checked connectivity, I can craft more checks of the DataStore methods themselves later.

The DataStore has a couple of methods that expose the Mongo::DB class that's underneath.  With that I can force a query for the list of collections currently available in the database. If the broker service has not yet been run and users and applications created then only the system collections will exist.  In the example below there are only two collections.


rails console
Loading production environment (Rails 3.0.13)
irb(main):001:0> store = OpenShift::DataStore.instance
=> #<OpenShift::MongoDataStore:0x7f1ac50ac698
 @host_port=["data1.example.com", 27017], @db="openshift",
 @user="openshift", @replica_set=false, @password="dbsecret",
 @collections={:application_template=>"template", :user=>"user", :district=>"district"};gt;
irb(main):002:0> collections = store.db.collections
=> [#<Mongo::Collection:0x7f1ac5096ac8 @cache_time=300, 
...
 @pk_factory=BSON::ObjectId>]
irb(main):003:0> collections.size
=> 2
irb(main):004:0> collections[0].name
=> "system.users"
irb(main):005:0> collections[1].name
=> "system.indexes"

On the MongoDB host I can confirm that there are indeed two collections.

mongo --username openshift --password dbsecret data1.example.com/openshift
MongoDB shell version: 2.0.2
connecting to: data1.example.com/openshift
> show collections
system.indexes
system.users

Finally I can check that the broker app really did issue that query and get a response:


...
Thu Dec  6 14:06:29 [conn2] Accessing: openshift for the first time
Thu Dec  6 14:06:29 [conn2]  authenticate: { authenticate: 1, user: "openshift",
 nonce: "d2083e4185cb7d22", key: "c7c3628fe64eb1aedaaf4c87a4d5e723" }
Thu Dec  6 14:06:29 [conn2] command openshift.$cmd command: { authenticate: 1, u
ser: "openshift", nonce: "d2083e4185cb7d22", key: "c7c3628fe64eb1aedaaf4c87a4d5e
723" } ntoreturn:1 reslen:37 5ms
Thu Dec  6 14:06:29 [conn2] query openshift.system.namespaces nreturned:3 reslen
:142 0ms
...

Checking Communications: Write


Now that I'm convinced that I'm connecting to the right database and I'm able to make queries, the next check is to be sure I can write to it when needed.

Since the database has not yet been used, it's empty.  I want to be careful regardless not to mess with any real OpenShift collections. I'll create a test collection, write a record to it, read it back and drop the collection again.  If I do this in a consistent way I can use this test at any time to check connectivity without danger to the service data.

 I'm using the ruby Mongo classes underneath the OpenShift::MongoDataStore class, so I'll have to look there for the syntax. The Mongo::DB class has a create_collection() method which will do the trick.  I'll issue the command in the rails console, then check the MongoDB logs and view the list of collections using the mongo CLI tool.

Create a Collection


First, the create query (entered into an existing rails console session):

irb(main):005:0> store.db.create_collection "testcollection"
=> #<Mongo::Collection:0x7f76567e9ae0 @cache_time=300,...
...
 @name="testcollection", @logger=nil, @pk_factory=BSON::ObjectId>
irb(main):006:0>

Next I'll check logs:

...
Thu Dec  6 14:37:13 [conn4] run command openshift.$cmd { authenticate: 1, user: 
"openshift", nonce: "665cedb4baf82b0d", key: "eec7b08761151c858c14058c2629dee6" 
}
Thu Dec  6 14:37:13 [conn4]  authenticate: { authenticate: 1, user: "openshift",
 nonce: "665cedb4baf82b0d", key: "eec7b08761151c858c14058c2629dee6" }
Thu Dec  6 14:37:13 [conn4] command openshift.$cmd command: { authenticate: 1, u
ser: "openshift", nonce: "665cedb4baf82b0d", key: "eec7b08761151c858c14058c2629d
ee6" } ntoreturn:1 reslen:37 0ms
Thu Dec  6 14:37:13 [conn4] query openshift.system.namespaces nreturned:3 reslen
:142 0ms
Thu Dec  6 14:37:13 [conn4] run command openshift.$cmd { create: "testcollection
" }
Thu Dec  6 14:37:13 [conn4] create collection openshift.testcollection { create:
 "testcollection" }
Thu Dec  6 14:37:13 [conn4] New namespace: openshift.testcollection
Thu Dec  6 14:37:13 [conn4] adding _id index for collection openshift.testcollec
tion
Thu Dec  6 14:37:13 [conn4] build index openshift.testcollection { _id: 1 }
Thu Dec  6 14:37:13 [conn4] external sort root: /var/lib/mongodb/_tmp/esort.1354
804633.1660751058/
Thu Dec  6 14:37:13 [conn4]   external sort used : 0 files  in 0 secs
Thu Dec  6 14:37:13 [conn4] New namespace: openshift.testcollection.$_id_
Thu Dec  6 14:37:13 [conn4]   done building bottom layer, going to commit
Thu Dec  6 14:37:13 [conn4]   fastBuildIndex dupsToDrop:0
Thu Dec  6 14:37:13 [conn4] build index done 0 records 0.001 secs
Thu Dec  6 14:37:13 [conn4] command openshift.$cmd command: { create: "testcolle
ction" } ntoreturn:1 reslen:37 1ms
...

Finally I'll connect and query the database locally to check for the presence of the new collection.

mongo --username openshift --password dbsecret data1.example.com/openshift
MongoDB shell version: 2.0.2
connecting to: data1.example.com/openshift
> show collections
system.indexes
system.users
testcollection

This is really enough to demonstrate that the MongoDataStore object is properly configured and has the ability to read and write the database.  Just for completeness I'll go one step further and create a document.

Add a Document to the testcollection


Since the testcollection is the most recently added, it should be the last one in the collections list in the Rails console Mongo::DB object.  I can check by looking at the name attribute of that collection

irb(main):007:0> store.db.collections[2].name
> "testcollection"

Now that I know I have the right one, I can add a document to it using the Mongo::Collection insert() method:

irb(main):008:0> store.db.collections[2].insert({'testdoc' => {'testkey' => 'testvalue'}})
=> BSON::ObjectId('50c0c2016892df2d56000001')

The logs show the insert like this:

Thu Dec  6 16:04:36 [conn11] run command openshift.$cmd { getnonce: 1 }
Thu Dec  6 16:04:36 [conn11] command openshift.$cmd command: { getnonce: 1 } nto
return:1 reslen:65 0ms
Thu Dec  6 16:04:36 [conn11] run command openshift.$cmd { authenticate: 1, user:
 "openshift", nonce: "19ca25a92ca483ee", key: "f7ac36d2e36a3a00a91d234a59a559e3"
 }
Thu Dec  6 16:04:36 [conn11]  authenticate: { authenticate: 1, user: "openshift"
, nonce: "19ca25a92ca483ee", key: "f7ac36d2e36a3a00a91d234a59a559e3" }
Thu Dec  6 16:04:36 [conn11] command openshift.$cmd command: { authenticate: 1, 
user: "openshift", nonce: "19ca25a92ca483ee", key: "f7ac36d2e36a3a00a91d234a59a5
59e3" } ntoreturn:1 reslen:37 0ms
Thu Dec  6 16:04:36 [conn11] query openshift.system.namespaces nreturned:5 resle
n:269 0ms
Thu Dec  6 16:04:36 [conn11] insert openshift.testcollection 0ms

And a quick CLI query to confirm that the document has been created:

mongo --username openshift --password dbsecret data1.example.com/openshift
MongoDB shell version: 2.0.2
connecting to: data1.example.com/openshift
> db.testcollection.find()
{ "_id" : ObjectId("50c0c2016892df2d56000001"), "testdoc" : { "testkey" : "testvalue" } }

Read a Document from the testcollection


In traditional database style, when you make a query, you don't get  back the single thing you asked for.  You get a Mongo::Cursor object which collects all of the documents which match your query.  Cursors respond to a next() method which does what you would think, returning each match in turn and nil when all documents have been retrieved.  The Mongo::Cursor also has a method which converts the entire response into an array. I'll use that to get just the one I want.

irb(main):035:0> store.db.collections[2].find.to_a[0]
=> #<BSON::OrderedHash:0x3fbb2b27fea4
 {"_id"=>BSON::ObjectId('50c0c2016892df2d56000001'),
 "testdoc"=>#<BSON::OrderedHash:0x3fbb2b27fd50 {"testkey"=>"testvalue"}>}>

I won't take up space showing the log entries for this query.  I know how to find them now if there's a problem.

Cleanup: Remove the testcollection


The final step in a test like this is always to remove any traces.  I can drop the whole collection with a single command.  This one I will confirm with the local CLI query, but the logs I'll leave for an exercise unless something goes wrong.

irb(main):037:0> store.db.collections[2].drop
=> true

You may notice that this was WAY too easy. Do be careful when you're working on production systems. Prepare and test backups OK?

When I look now on the CLI and ask for the list of collections, I only see two:

mongo --username openshift --password dbsecret data1.example.com/openshift
MongoDB shell version: 2.0.2
connecting to: data1.example.com/openshift
> show collections
system.indexes
system.users

Summary

In this post I showed how to access the OpenShift broker application using the Rails console.  I created an OpenShift::MongoDataStore object (using the OpenShift::DataStore factory).  I showed how to access the database from the CLI and where to find the MongoDB log files.  With these I was able to confirm that the OpenShift broker DataStore configuration was correct and that the database was operational.

References






Friday, November 16, 2012

OpenShift Back End Services: Data Store (MongoDB)

In the previous post, I listed the set of back end services which underpin an OpenShift Origin service.
In this one I'm going to detail the installation of the first of these back end services and initialize it for OpenShift Origin.

Data Persistance

The OpenShift Origin service needs a backing data store to contain persistent data.  It must keep track of the meta data for user's applications, the available nodes and their capabilities and the ssh public keys provided by the user to allow them git and ssh access to their apps.

Currently OpenShift Origin uses a MongoDB back end for data persistence.  Before you can start building your OpenShift Origin service you need to have a running MongoDB that can be reached by the broker service on the Openshift Origin broker host(s).

There are lots of good resources about administration and use of MongoDB.  If you're going to run an OpenShift Origin service yourself you should keep them handy.  Check the References section at the end of this (and each) post.

In this post I'm going to walk through preparing the Data Storage host.  Except for a couple of commands to create the user accounts and an empty database, there isn't anything really databasey about this procedure.

Information Gathering

From the table in the last post, we have the hostname and IP address of the broker host and the MongoDB server.

We're going to install and configure the MongoDB service to permit the OpenShift Origin broker service to access, read and write the OpenShift database.

The MongoDB service runs on port TCP 27017.

I'm going to create a root account in the MongoDB admin database and enable authentication.  This will allow two layer access control to the data.  I'll create the empty OpenShift Origin database so that it is ready for the broker to connect. I will also create a role user in the OpenShift database for the Openshift Origin  broker to use.  I'm going to create values for those.  You should choose different passwords.

I will also need the IP address of the broker host so that I can use the iptables firewall to limit inbound connections to only that host.

So here's the information we have:

MongoDB Setup Information
FunctionHostnameIP Address
Broker Hostbroker1.example.com192.168.5.11
Data Storage Hostdata1.example.com192.168.5.8
MongoDB TCP Port27017
OpenShift Database Nameopenshift
AccountUsernamePassword
MongoDB Privileged Userrootdbadminsecret
Openshift Database Useropenshiftdbsecret
 

Preparing the Base

In each of these setups I assume that the hosts have a base operating system installed and configured.  I'm working with RHEL 6.3 and Fedora 17, but some of the packages are not yet publicly available for those. (November 15 2012) They will be available with the release of Fedora 18 at the end of the month and either from EPEL or from RHEL repositories.

I start with the Base package list, ntpd, policycoreutils-python and what ever packages are needed for central user control.  These are Kerberos 5 and LDAP in my case (krb5-workstation, pam_krb5 and openldap-clients).  For me this works out at between 300 and 425 packages. Right now I have to add a set of yum repositories for the pre-release builds but these should not be neccesary once all of the packages go to release for both RHEL and Fedora.

I also generally bring my base system up to date before beginning any other work.

yum -y update

The Process

There are several steps to preparing MongoDB for the OpenShift Origin service. I have to make sure it will start correctly, that it is secured and that it is initialized so that when the broker first connects the database is ready to accept updates.  I also need to save the information that the broker will need to establish a connection as I'll need that to configure the broker plugin.

The steps look like this:
  • Install the MongoDB server software
  • Enable authentication
  • Add an administrator account
  • Add the empty OpensShift database
  • Add the OpenShift broker user account
  • Add firewall rules
  • Enable listening on external interfaces
  • Enable service restart on boot
Each of these steps is fairly small and the whole thing should be easily scriptable. It also should be fairly simple to script a set of checks to verify that the database service is properly configured and secured. These scripts can be used to check the underlying database service configuration in the event of a problem with the OpenShift service.

Note that I try to perform the steps in such a way that the service is never exposed to an external network until authentication is enable and configure and the firewalls have been established.  This prevents cracking attempts during the (admittedly tiny) window when anonymous access would succeed.

Installing the Software

The first thing to do is to install the mongodb-server software. This will pull in several pre-requisites as well.

yum -y install mongodb-server

Enable Authentication

The MongoDB configuration file is /etc/mongodb.conf. The configuration is a simple space/line delimited key/value pair format. Authentication is off by default. The auth section looks like this:


...
# Turn on/off security.  Off is currently the default
#noauth = true
#auth = true
...

The first thing to do is make a copy of the original configuration file in case I mess up.
I need to uncomment the "auth = ..." line. You can do this with an editor, but I generally do this for simple changes with a line editor like sed(1). I'm also going to sneak in one other tuning parameter here that OpenShift wants but isn't related to security or service management: The smallfiles parameter needs to be added to the end of the file and set to true.
cp /etc/mongodb.conf /etc/mongodb.conf.orig
sed -i -e '/^#auth =/auth =/' /etc/mongodb.conf
echo "smallfiles = true" >> /etc/mongodb.conf

Throughout these pages you'll see patterns like that: "save a copy, make a change or two".

Add Accounts and Empty Database.

To add accounts and initialize the database it must be running.  Once it is, I'll use the mongo(1) CLI client tool to make the changes I need.
The mongodb process needs a few seconds to start and begin listening, so there's a sleep after starting the daemon and before trying to access the database.
There are three steps here:

  • Create the administrator (root) account.
  • Create the OpenShift Origin server database
  • Create the OpenShift Origin broker role account.
MongoDB will actually create a database just by being told to use it even if it doesn't exist, so the steps are actually simpler. I'm also going to stop the database service once this is done so that I can open it (carefully) to external network access.

Note that the passwords and the OpenShift Origin database name come from the table of values a the beginning of this post. You need to change the password and you can select a different database name if you wish (you must keep track of it, you'll need it later).

# Start the mongod service (obviously)
service mongod start

# Wait for the service to be ready to listen
sleep 10

# Connect, create root account, authenticate, create database and role account.
mongo admin << EOMONGO
db.addUser('root', 'dbadminsecret');
db.auth('root', 'dbadminsecret');
use openshift;
db.addUser('openshift', 'dbsecret');
EOMONGO

# Stop the service again.
service mongod stop

In the past I would have also created a read-only user in the admin database. This would have allowed full db backups without read-write access. Now I should probably create a read-only user in the openshift database to back up just the one database.

Open a small hole in the firewall

By default the mongod process only listens on the localhost interface. Since it also starts with no authentication this is a good thing for security. Now I want to allow the service to listen for outside connections. However I only want to allow connections from the OpenShift Origin broker host. I'll use an iptables entry to restrict inbound connections and then tell the daemon to listen on the external interface.

If you wanted to allow unrestricted inbound connections, you could just use lokkit but we want to restrict access to a single or small set of inbound hosts so we need to configure the IP tables deliberately.

It would be handy to have the iptables(8) man page handy here for reference.

I want to allow inbound connections only from the broker host. I'll have to craft an iptables line which will do that. I have the IP address of the broker host and the TCP port to which mongod listens by default. I'll use those to craft the appropriate rule.

-A INPUT-s 192.168.5.11/32 -m state --state NEW -m tcp -p tcp --dport 27017 -j ACCEPT

This line means:

"Append an entry to the INPUT queue. Match NEW connections from 192.168.5.11. Accept TCP packets destined for port 27017"

I'd like to add this line to the end of the INPUT queue, but actually not the last line. The last line is the one that says "anything not matched by now, reject it". I want that to stay at the end. I want to insert my allow line just before that.

I could craft a clever sed or awk command to line edit the /etc/sysconfig/iptables file. iptables itself provides me with a better way. I can use the iptables control commands to determine which rule I want to insert at, add the line to the running tables and then have iptables dump the result to a file. I can save that as the new /etc/sysconfig/iptables file so that it will restore my new rule set at system startup.

I can use the iptables command first to list the existing ruleset. I can count the number of rules with wc -l and use expr to subtract one. That becomes the rule number for my insert command.

# Save a copy of the original
cp /etc/sysconfig/iptables /etc/sysconfig/iptables.orig

# Find the index of the next-to-last rule in the INPUT ruleset
INSERT_INDEX=$(expr $(iptables -S INPUT | wc -l) - 1)

# Add a new rule at N-1
iptables -I INPUT $INSERT_INDEX -s 192.168.5.11/32 -m state --state NEW -m tcp -p tcp --dport 27017 -j ACCEPT

# dump a copy of the current rules and save them for next reboot
iptables-save > /etc/sysconfig/iptables

Listen on external interfaces

Now that the authentication and firewall are in place, I can reconfigure mongod to listen on all interfaces. We're back to the easy stuff.

The bind_ip option in the /etc/mongodb.conf file specifies where the mongod will bind. By default this is 127.0.0.1. The IPv4 convention for "all addresses" is 0.0.0.0. I'll replace the value of bind_ip in /etc/mongodb.conf and I'll be ready to restart the service. This is a simple sed script again.

sed -i -e '/^bind_ip =/s/=.*/= 0.0.0.0/' /etc/mongodb.conf
service mongod start

Access and Security Verification

Once the service is running I want to make attempts to connect to it from at least three different sources.
  1. From the datastore host (localhost) - allowed
  2. From the datastore host (data1.example.com) - rejected
  3. From the broker host - allowed
  4. From an unauthorized host - rejected
I also want to connect both to the admin database and to the openshift on the two that should not be restricted by the firewall.

Note that, because the firewall rule did not include an allow clause for localhost as the source, you can only connect to the database using the localhost interface if you are on the datastore host.

The test commands below consist of an echo command which writes a string to standard output. The string is a single mongodb CLI command.  The string is piped as input to the mongo CLI client.  The argument to the mongo command is the host and database to be opened. 

# admin, root user on localhost
echo 'db.auth("root", "dbadminsecret");' | mongo localhost/admin
MongoDB shell version: 2.2.1
connecting to: localhost/admin
1
bye

# openshift, openshift user on localhost
echo 'db.auth("openshift", "dbsecret");' | mongo localhost/openshift
MongoDB shell version: 2.2.1
connecting to: localhost/openshift
1
bye

# admin, root user on external interface
echo 'db.auth("root", "dbadminsecret");' | mongo data1.example.com/admin
MongoDB shell version: 2.2.1
connecting to: data1.example.com/admin
Fri Nov 16 21:28:45 Error: couldn't connect to server data1.example.com:27017 src/mongo/shell/mongo.js:93
exception: connect failed

# openshift, openshift user on external interface.
echo 'db.auth("openshift", "dbsecret");' | mongo data1.example.com/openshift
MongoDB shell version: 2.2.1
connecting to: data1.example.com/openshift
Fri Nov 16 21:28:45 Error: couldn't connect to server data1.example.com:27017 src/mongo/shell/mongo.js:93
exception: connect failed
I would try those two tests on the broker host and some external host to verify access. I should probably add a test to each with invalid user and with valid user but invalid password to be rigorous.

Don't Do This (Without Re-Compiling with SSL)

Having gone through this exercise to try to show how the MongoDB back end can be disentangled from the Openshift Origin service configuration, I have to say: Don't Do It.

The last step should have been to establish an encrypted pipe between the broker host and the datastore database.  It turns out that MongoDB and most NoSQL databases don't think they should be doing encryption, so they don't.  While they taught the benefits of distribution and sharding and other multi-host based behaviors, I really can't recommend allowing MongoDB to play in the street with the other kids.  Both the authentication information and the data go in clear text.

In the MongoDB: The Definitive Guide  from O'Reilly the authors recommend an SSH tunnel if you really need encryption point to point.  I don't find this very satisfying.

The MongoDB documentation web site has a page on using MongoDB with SSL.  It requires recompiling the package with a -ssl option.  The documentation does not give a detailed set of instructions for re-compiling.  You have to work that out for yourself.

If you install the resulting package you then have to mark your yum repository so that it doesn't update the mongodb RPM.  Updates from the stock repositories will destroy your SSL configuration.  When an update is indicated, you'll have to download the new source RPM, rebuild again and manually update the package.

For now, if you don't recompile with SSL I'd suggest that OpenShift servers be restricted to a single broker with a single unreplicated database on a single host.

The exercise remains a good one.  You can follow the steps listed here using localhost for all of the datastore host addresses and you'll see the parts of the setup that are specific to the datastore as distinct from the OpenShift Origin service proper.

Next Up

Next up will be the Messaging service using ActiveMQ.

References