Implementing an S3 native filesystem in Hadoop
Let's first create
OutputStream for the filesystem. In our example, we have to connect to the AWS to read and write files to S3.
Hadoop provides us with the
FSInputStream class to cater to custom filesystems. We extend this class and override a few methods in the example implementation. A lot of private variables are declared along with the constructor and helper methods to initialize the client as illustrated in the following code snippet. The
private variables contain objects that are used to configure and retrieve data from the filesystem. In this example, we use objects such as
AmazonS3Client to call REST web APIs on AWS,
S3Object as a representation of the remote object on S3, ...