Thursday, March 26, 2009

The PrimeBase BLOB Streaming (PBMS) engine alpha version 5.08 is ready

Alpha version 5.08 of the BLOB streaming engine for MySQL has been released. You can download the source code from www.blobstreaming.org/download. The documentation has also been updated.

What's new in 5.08:
  • All PBMS data is stored under a 'pbms' directory in the MySQL server's data directory rather than in the database directories them selves.
  • This version now builds with Drizzle and can be loaded as a 'Blobcontainer' plug-in.
  • Added the possibility of storing BLOB metadata along with the BLOB in the repository.
  • Added the possibility of assigning an alias to a BLOB, which can then be used to retrieve the BLOB instead of using the engine generated URL.
  • Added an updateable system table 'pbms_metadata_header' to control which HTTP headers are stored as metadata.
  • Added an updateable system table 'pbms_metadata' that contains all the metadata associated with the BLOBs.
  • New PBMS API functions have been added to set and get BLOB metadata.
  • A new PBMS API function has been added to allow applications to get the BLOB metadata with out getting the actual BLOB.
  • Added some new fields to the 'pbms_repository' system table.
  • Removed the raw BLOB data from the 'pbms_repository' system table and placed it in its own table, 'pbms_blob'.
  • Dropping a database containing PBMS BLOBS referenced from non-PBXT tables no longer requires special handling.
  • System tables can now be selected with an 'order by' clause.
As you can see a lot of work has been done on it and a lot of things have changed. I have already talked about most of the major changes in my previous 2 BLOG postings but I recommend having a look at the documentation for more details.

A few things to watch for in the new version:

  • The location and format of the BLOB repository files has changed. This means that if you are using an older version you will need to import the data from the server using the older version of PBMS into tables on a server running the new version. Feel free to contact me if you have any questions about this. I will try to maintain backward compatibility with older versions of PBMS but until the code is no longer alpha I can not guarantee this.
  • A patch was added to the PBXT engine to prevent a crash when shutting down MySQL. The PBXT version '1.0.07m-rc' contains this patch and is available for download with the the PBMS engine.
  • There remains an unsolved problem that can lead to PBXT hanging when a PBXT table containing longblobs is dropped and the PBMS engine is being used for BLOB storage.
As usual if you have any questions about PBMS and BLOB storage please send them to me and I will do my best to answer them.

Barry

Wednesday, March 4, 2009

PBMS supports BLOB metadata and aliases

The PrimeBase PBMS engine now supports user defined metadata.

When the PBMS engine receives a BLOB the HTTP header tag value pairs are stored with the BLOB as metadata. To restrict which headers are stored as metadata an updatable system table 'pbms_http_fields' is provided in which users can add the names of the headers that are to be treated as metadata. When the BLOB is retrieved from the engine the metadata is sent back along with the BLOB as HTTP headers in the same way that it was received.

The metadata can be accessed from within the database via the 'pbms_matadata' system table or altered by performing inserts, updates, or deletes on the table.

A BLOB alias is a special metadata field that allows you to associate a database wide unique name with the BLOB which can then be used later to retrieve it. If you are familiar with Amazon S3 storage it works in a similar manner where you can think of the database as the S3 bucket and the BLOB alias as the S3 key.  To fetch the BLOB back using the alias you use a URL with the format <database>/<alias>. The following is an example using 'curl' to send a BLOB with alias 'MyBLOB' to the PBMS engine and then fetch it back again:

curl -H "PBMS_BLOB_ALIAS:MyBLOB" -d "A BLOB with alias" "http://localhost:8080/test"
Returning:
~*test/_1-632-4934f86e-0*17

curl -D - "http://localhost:8080/test/MyBLOB"
Returning:
HTTP/1.1 200 OK
PBMS_CHECKSUM: A3762FF16159FAB246EBA2BE50F98CF4
PBMS_BLOB_SIZE: 17
PBMS_LAST_ACCESS: 1236200654
PBMS_ACCESS_COUNT: 0
PBMS_CREATION_TIME: 1236200654
PBMS_BLOB_ALIAS: MyBLOB
Content-Length: 17

A BLOB with alias
In this example the BLOB has been uploaded to the PBMS engine but not yet referenced so it will be automatically deleted after a preset time. 

As you can see several custom headers have been added to the reply:
  • PBMS_CHECKSUM is the MD5 checksum of the BLOB data.
  • PBMS_BLOB_SIZE is the size of the BLOB data in bytes.
  • PBMS_LAST_ACCESS is the last access time of the BLOB in seconds since Jan. 1 1970.
  • PBMS_ACCESS_COUNT is the number of times the BLOB has been downloaded.
  • PBMS_CREATION_TIME is the creation time of the BLOB in seconds since Jan. 1 1970.
It is also possible to fetch back just the header info with out the BLOB data, example:
curl -H "PBMS_RETURN_INFO_ONLY:yes" -D - "http://localhost:8080/test/MyBLOB"
Returning:
HTTP/1.1 200 OK
PBMS_CHECKSUM: A3762FF16159FAB246EBA2BE50F98CF4
PBMS_BLOB_SIZE: 17
PBMS_LAST_ACCESS: 1236200654
PBMS_ACCESS_COUNT: 0
PBMS_CREATION_TIME: 1236200654
PBMS_BLOB_ALIAS: MyBLOB
Content-Length: 0
The PBMS BLOB metadat enables users to store BLOB specific data with the BLOB and have it returned to them with out having to create a separate database table to store it and then execute separate SQL command to update it and retrieve it when ever they upload or download a BLOB. 

The BLOB alias allows users to generate their own names for BLOBs which they can then use to access the BLOB without having to make a call to the database to get the PBMS generated URL.



Thursday, February 26, 2009

PBMS supports Drizzle

The PBMS engine now works with Drizzle. Well actually it has been working with Drizzle for several months since I have been using Drizzle as my 'host' server while adding new features to the engine. I will tell you about the new features in future posts.

Hooks for a 'blobcontainer' type plug-in have been added to Drizzle that allow a plug-in to catch insert. update, or delete operations on BLOB columns in any table and handle the BLOB storage itself. The plug-in gets called above the storage engine level so it is independent of the storage engines. This is very similar to the way that PBMS works in MySQL 5.1 with engines other than PBXT where it uses triggers to perform the same function. But the way it is done with Drizzle is a lot more efficient.

So in Drizzle the PBMS engine is both a 'blobcontainer' plug-in as well as a storage engine. 

Surprisingly little needed to be done to the PBMS engine to get it to build with Drizzle. It was a big help having the PBXT engine already ported to Drizzle so for a large part all I needed to do was look to see what Paul had done and do the same. The same source code is used for both the Drizzle and MySQL builds and the build procedure is the same. When you configure the engine different build flags will automatically be set depending if the target is a MySQL source tree or a Drizzle source tree. But before you try doing a build yourself you should know that the 'blobcontainer' plug-in support has not been merged into the main Drizzle build branch yet. 

The one drawback to the way the 'blobcontainer' plug-in works currently is that it take a shotgun  approach to BLOB management in that either all blobs in all tables are stored using the 'blobcontainer' plug-in or none are. This is because there is no way of telling the plug-in which BLOB columns are to be stored using the plug-in and which are not. I am in hopes that some way will be found so that BLOB columns can be selectively handled by the 'blobcontainer' plug-in. Possible ways of doing this would be to have a special BLOB column type for BLOBs to be stored with the 'blobcontainer' plug-in or provide a column attribute to signal how the BLOB data is to be stored. In MySQL 5.1 only longblob columns are stored using the PBMS engine. I have no doubt though that a good solution for this problem will be found before long.

Barry

Wednesday, November 5, 2008

Amazon S3 file transfer daemon

This is a project that I have been working on that I have just uploaded to launchpad. It is not directly related to my work on the BLOB streaming engine for MySQL but it does contain code and ideas that I hope to work into it. The project name on launchpad is 's3daemon'.

The Amazon S3 file transfer daemon is a daemon process that runs in the background and monitors folders. When it finds a file in one of the folders that matches a specific pattern it transfers that file to S3 storage and then removes the local copy. The folders monitored and the patterns searched for are controlled by a configuration file.

Included with the daemon is an apache module that handles requests for files in the folders monitored by the daemon. If a request comes in for a file and the file cannot be found locally the caller is sent a redirect to get the file directly from the S3 server.  The redirect contains a signature created by the module using a public/private key combination that tells the S3 server that the caller is authorized to get the requested file. Included in the authorization is a time stamp that limits how long the URL/signature combination is good for.

I am thinking that a BLOB streaming engine that handled BLOBs in a similar manner may be of interest to people storing a large amount of image data. Such an engine would allow the creation and deletion of the data to be controlled by the database while the actual BLOBs are stored in S3 storage. This also decreases the bandwidth requirements of the server machine because the actual data will be served up by the S3 server.



Monday, October 20, 2008

Using config.status to build outside the tree

I just thought I would share some of my discoveries. You may already have known this but I didn't.

In recent releases both the PBMS and PBXT engines have been using a configuration flag "--with-mysql" to tell the build where to find the MySQL tree. We then looked inside the 'config.status' file generated when MySQL was configured to get the build options, most importantly the compiler options used. We had been doing this using 'grep' and 'sed'. The problem we soon discovered was that the format of the 'config.status' file changes quit frequently with different versions of 'autoconf'. 

With the help of the good people on the 'autoconf' forum I discovered that you can ask 'config.status' to give you the value of a single substitution. So to get the value for CFLAGS you enter:
echo '@CFLAGS@' | config.status --file-
and it will print out a line with the CFLAGS.

I understand that this feature which has always existed is now being added to the 'autoconf' documentation.

Using this when configuring  your build guarantees that your compiler options match that of MySQL and avoids those nasty bugs where structure alignments do not match because of different compiler settings.

I hope this helps someone out there.

Barry

Tuesday, October 14, 2008

Ideas for BLOB streaming

Hi,

How that I have the latest release of PBMS out the door I thought I would post some of the ideas I have had for possible ways in which the BLOB streaming engine could be expanded upon. 

What if the BLOB streaming engine supported the idea of having it's own BLOB storage engines that could be plugged into it the same way that storage engines are plugged into MySQL. The API for these BLOB storage engines would be dead simple, all they would need to support would be a 'get', 'put', and 'delete' method. The BLOB streaming engine would handle the reference counting and still provide the simple HTTP server for direct access to the BLOBs but the BLOB engines would handle how and where the actual BLOB data would be stored.

Currently the blob data is stored locally in blob repository files, which is very efficient but may not be ideal for some applications. Here are a few ideas I have had for possible BLOB storage engines:
  • File Storage: This storage engine would just store the blobs in individual files. It may be of interest to applications where they want to be able to directly grab the data such a web application where the data is directly read by the web server.
  • Amazon S3 Storage: This storage engine would store the BLOBs in Amazon S3 buckets. This BLOB storage engine would be of interest to applications that deal with massive amounts of data and need a highly scalable storage solution.
  • Mirrored Storage: This storage engine would replicate the BLOBs to multiple geographical locations so that requests for the data could be directed to the closest mirrored site.
  • Load Balancing Storage: This storage engine would have the ability to move the BLOB data around to from one storage location to another to try to optimize it's access time or more evenly distribute the storage load on different servers.
  • Wally Storage: This engine wouldn't actually store anything and would reply to any request for data that it will have it for them by Wednesday. This would be of interest to applications that are forced to store data that they know nobody will ever actually want.
To allow for more flexibility in the BLOB storage engines the BLOB streaming engine would have to be able to handle a redirect reply from the BLOB storage engines. This redirect could be handled inside of the PBMS client API so that an application requesting a BLOB via the API would never need to know anything about the redirect that took place.

This concept would allow the manner and location in which BLOBs are stored to be completely decoupled from the database and database server.


Barry

Alpha release v05.06 of the BLOB streaming engine


Alpha version 5.06 of the BLOB streaming engine for MySQL has been released. You can download the source code from www.blobstreaming.org/download. The documentation has also been updated.

What's new in 5.06:
  • The BLOB streaming engine can now be used with MyISAM tables as well as tables created by any other MySQL storage engine. 
  • The name of the PrimeBase BLOB streaming engine has been changed from MyBS to PBMS which stands for "PrimeBase Media Streaming".
This version introduces a couple of new term:
  • streaming-enabled table: This is any table created by a streaming-enabled engine such as PBXT, or has triggers defined on blob referencing columns that notify the PBMS engine when BLOB references are inserted, updated, and deleted.
  • BLOB reference column: This is a column in a stream-enabled table that contains PBMS BLOB references and notifies the PBMS engine when data is inserted, updated, or deleted from the column.
The big news though is that as of this version you no longer need to have PBXT installed in order to use the PBMS engine. The PBMS engine provides a set of UDFs and client API functions that enable it to be used with tables created by any storage engine. The following is an example of how to create a streaming enabled MyISAM table:

Create table x.foo(c1 integer, c2 longblob) ENGINE = MYISAM;

create trigger x.foo_insert_trig BEFORE INSERT on x.foo for each row BEGIN set NEW.c2 = pbms_insert_blob_trig("x", "foo", 2, NEW.c2); END

create trigger x.foo_update_trig BEFORE UPDATE on x.foo for each row BEGIN set NEW.c2 = pbms_update_blob_trig("x", "foo", 2, OLD.c2, NEW.c2); END

create trigger x.foo_delete_trig BEFORE UPDATE on x.foo for each row BEGIN declare dummy integer; set dummy = pbms_delete_blob_trig("x", "foo", 2, OLD.c2); END


As a result of the name change you will see that any use of the letters 'mybs' has been changed to 'pbms' throughout the code and documentation.

Of lesser importance but also of note: I have become the official contact person for the BLOB streaming engine. Paul is still very much involved with it but is currently concentrating his efforts on the PBXT engine.

Barry